INTEL (NL)
nl

How We Stopped Relying on JavaScript and Finally Got LLMs to Read Our Website Data

Ontdek hoe AI-zoekmachines webpagina's scrapen. Leer waarom JavaScript faalt en hoe Markdown en llms.txt uw vindbaarheid in LLM's garanderen.

AnswerShaper Editorial
16/08/2026
11 min. leestijd
How We Stopped Relying on JavaScript and Finally Got LLMs to Read Our Website Data

How We Stopped Relying on JavaScript and Finally Got LLMs to Read Our Website Data

The $50,000 JavaScript Mistake That Made Us Invisible

Watching our traffic die in real-time

I spent 3 hours last night testing our URLs in different LLMs.

It was brutal.

Six months ago, we dropped $50,000 on a massive website overhaul. We wanted the wow factor. We built a gorgeous, animation-heavy, JavaScript-rendered masterpiece. It had dynamic scroll effects, interactive pricing sliders, and lazy-loaded case studies. It looked like a million bucks to anyone browsing on a MacBook.

But nobody was browsing.

Staring at our analytics dashboard at 2 AM, the reality hit me hard. Inbound leads had dropped by 40% over three weeks. The organic traffic line looked like a cliff. The hard truth? Buyers aren't clicking on your site anymore. They are asking AI.

So, I ran an experiment. I opened ChatGPT, Claude, and Perplexity. I pasted our shiny new URLs into the prompt boxes and asked a simple question: What does this company do, and how much does it cost?

Total, absolute hallucination.

ChatGPT confidently invented a $99/month starter tier that we never offered. Claude completely missed our flagship enterprise product. Perplexity scraped a broken, cached version of a footer from two years ago.

Why? Because they couldn't read our site.

Our $50,000 investment was entirely rendered in client-side JavaScript. When the LLM crawlers hit our domain, they didn't wait for the beautiful animations to load. They didn't interact with the pricing sliders. They hit a blank HTML shell, saw a massive <script> tag, hit their token limit, and bounced. Because they couldn't execute the JS, they just guessed the rest based on outdated training data.

The vrai problème isn't that AI is hallucinating. The real problem is that we are still building websites for human eyes instead of machine parsers.

We built a digital billboard in a world where machines are doing the reading. We spent thousands optimizing for human attention, completely ignoring the bots that actually control the flow of information in 2026. If the algorithms can't parse your raw data instantly, they move on.

We were invisible.


Why Traditional Crawling is Dead in the M2M Era

That catastrophic drop forced us to ask a question we'd always dodged:

Can LLMs read HTML?

Yes, Large Language Models can read HTML, but they primarily rely on visible text, structured data, and metadata rather than executing complex scripts, meaning that heavily nested code or JavaScript-rendered elements often get ignored due to strict token limits and high computational costs during the parsing process.

Knowing that technical limitation made me furious. I am absolutely marre des conseils from so-called experts.

You know the ones. The gurus on your feed telling you to "just write good content" and let the search engines do the rest. That advice is a relic. It belongs in 2015.

LLMs are not Googlebot.

Googlebot has the time, the infrastructure, and the crawling budget to sit there and render your heavy JavaScript animations. AI agents do not. They face strict API costs and brutal token limit constraints. Every single nested <div> and inline script eats into their context window. If a machine has to execute complex JavaScript just to read your core product features, it simply won't. It will bounce, hallucinate an answer, and move on to your competitor with a cleaner DOM.

Relying on traditional SEO tactics for AI visibility is a guaranteed path to obsolescence.

If they aren't crawling our heavy JavaScript, we had to figure out their actual source.

Where does the data for LLMs come from?

The data for Large Language Models comes primarily from massive, pre-existing training datasets scraped from the public internet, licensed databases, and proprietary archives, meaning these models heavily prefer relying on this static historical data over performing live web fetching unless a user explicitly prompts them to browse.

This reliance on static training data is the vrai problème most marketing teams refuse to accept.

Founders think ChatGPT is out there actively crawling their newly published blog posts every morning. It isn't. Unless a user explicitly forces the model to fetch a live URL, the AI defaults to its training weights. It takes the path of least resistance.

Si votre SEO ne prend pas en compte le M2M (Machine-to-Machine), you are already losing.

We aren't optimizing for human eyes first anymore. We are optimizing for parsers, scrapers, and RAG pipelines. If your data isn't structured to be instantly digestible by a machine, it doesn't matter how pretty your website looks. The buyers aren't clicking anymore; they are asking AI. And if the AI can't read your site cheaply and efficiently, you simply don't make it into the output.


The Markdown Loophole That Changed Everything

We knew we weren't making it into the output. So, I decided to run a desperate experiment.

Stripping the web down to its bones

I was staring at VS Code, endlessly scrolling through the source code of our flagship landing page. It was a complete disaster. Raw HTML is a nightmare for machine parsers. It's packed with nested <div> tags, inline styling, tracking scripts, and noisy developer comments.

Every single one of those characters eats up precious context window tokens.

When an LLM tries to read that bloated mess, it chokes. It gets distracted by your hamburger menus and your legal footers. It burns through its token limit just trying to figure out where the actual article starts.

I'd had enough. I grabbed the text editor.

I deleted everything.

No CSS. No JavaScript. No HTML tags. I stripped the entire complex page down to a simple, bare-bones .md file. Just plain text, a few hash symbols for headers, and standard line breaks.

I fed that raw Markdown file directly into ChatGPT.

The result? Instant clarity.

By converting our core pages to clean Markdown, the AI stopped hallucinating. Instead of inventing 4 out of 5 of our pricing tiers, it extracted all 5 perfectly.

It wasn't just reading the page anymore. It actually understood the hierarchy. It grasped the relationships between our features without hallucinating a single pricing tier.

Here's the vrai problème with modern web design: we build for human eyes, not for neural networks. But LLMs parse Markdown significantly better than raw HTML.

Markdown is their native language.

It's clean. It's structured. It's exactly what the algorithms want to consume. When you use Markdown, you spoon-feed the context directly to the parser. An H2 is immediately recognized as a major section. A bolded term is flagged as a priority entity.

You stop wasting tokens on visual formatting. You start spending them on actual meaning.

Think about how these models are trained. They ingest massive repositories of GitHub documentation, Reddit threads, and plain text datasets. They are literally wired to understand Markdown natively. When we force them to crawl a heavy, modern web page, we're making them translate a foreign language before they can even read the message.

That late-night experiment changed our entire technical roadmap.

We realized we didn't need to build more complex scrapers or heavier SEO plugins. We just needed to speak the machine's language. We stopped fighting the algorithms and gave them exactly what they wanted.


The 3-Step Blueprint to Feed the Algorithms

We gave them what they wanted on one page, but we still had to scale it. We had to ask:

Can Chat GPT extract data from a website?

ChatGPT can extract data from a website if it has active browsing capabilities enabled, but it primarily relies on visible HTML content, structured data, and metadata while struggling heavily with client-side JavaScript rendering, complex animations, and strict token limits that prevent full-page crawling.

That is the baseline reality. But knowing that wasn't enough to fix our dropping traffic. We had to completely rewire our technical stack.

I was staring at our server logs, realizing bots were hitting our site and leaving with empty payloads. It was a vrai problème. We were serving them a blank canvas.

We needed a system. A ruthless, machine-first architecture.

Deploying llm.txt and Server-Side Rendering

We tore our infrastructure down to the studs. I'm marre des conseils from traditional SEOs telling you to just "write better content." Good content is useless if the machine can't parse it. Here is the exact three-step framework we deployed to force LLMs to read our data.

Step 1: Switch to Server-Side Rendering (SSR)

We ripped out our client-side rendering. Gone. We switched entirely to Server-Side Rendering (SSR) using Next.js.

Why? Because AI agents are cheap. They don't want to spend compute power executing your bloated JavaScript. If your pricing, features, and core value propositions aren't sitting right there in the raw HTML payload, the bot moves on. We watched our server response times drop, but more importantly, the raw source code finally contained actual words instead of empty script tags.

If the bot has to render JS to see your product, your product doesn't exist.

Step 2: Implement strict Schema.org and semantic HTML

Next, we killed the <div> soup.

I spent an entire weekend auditing every core page. We implemented strict Schema.org markup and semantic HTML tags—<article>, <section>, <aside>—to spoon-feed entities directly to the parser.

You can't expect an LLM to guess what your page is about. You have to map it out. We structured our FAQs with FAQPage schema. We tagged our product features with exact Product attributes. By stripping out the visual noise, we reduced our DOM size by 40%—dropping from over 3,500 DOM nodes to just under 2,100. That directly preserved precious token limits for the crawlers, ensuring they ingested our entire value prop before hitting their cutoff.

Step 3: Create a dedicated 'llm.txt' file

This was the final piece.

We added a dedicated llm.txt file straight to the root directory of our site. Think of it like a robots.txt, but instead of telling bots where not to go, it hands them exactly what they need.

This file provides a clean, Markdown-optimized version of our entire data structure directly to AI agents. No CSS. No tracking scripts. Just pure, unadulterated text formatted in the exact syntax LLMs natively understand. We linked it in our header. We pointed every custom GPT and agent directly to it.

We stopped hoping the algorithms would figure us out. We started feeding them directly.


Either You're in the Prompt, or You Don't Exist

Feeding them directly isn't just a tactic anymore. It's survival.

The brutal reality of 2026 search

The web broke.

We spent two decades obsessing over human eyeballs. We tracked heatmaps, optimized button colors, and wrote clever meta descriptions. None of that matters now.

Here is the cold truth: les acheteurs ne cliquent plus sur votre site. Buyers aren't clicking on your site anymore. They ask an AI agent, get their answer in three seconds, and make a purchasing decision without ever loading your CSS.

If you aren't structuring data for AI, you are invisible to the modern buyer.

Stop fighting the algorithms. Stop trying to force a 2015 SEO strategy onto a 2026 machine-to-machine reality. The bots don't want your beautiful animations. They want raw, structured, semantic data. They want it formatted exactly how they consume it.

You are no longer marketing to humans first. You are marketing to the parsers that advise the humans.

I see founders panic when their organic traffic flatlines. They stare at their Stripe dashboards, wondering where the inbound leads went. The buyers didn't leave. The interface just changed.

We got tired of manually checking this, so we use AnswerShaper. We needed a gritty utility to ping our URLs and tell us exactly what the parser sees before our competitors beat us to the context window. It strips away the noise and shows us if the machine actually understands our pricing, features, and core value proposition.

Adapt or disappear.

It really is that simple. You can keep writing for human eyes that will never see your page, or you can feed the machines exactly what they demand.

Soit vous êtes dans le prompt, soit vous n'existez pas.

Hoe LLM's Data Lezen: JavaScript en Markdown | AnswerShaper | AnswerShaper Blog