How a search engine actually works under the hood — crawl, index, rank, and the lies in between

Elena Varga

Elena Varga

September 18, 2026

How a search engine actually works under the hood — crawl, index, rank, and the lies in between

People say “search is just keywords” the way people say “the database is just tables.” I have built two tiny engines — one for a docs site, one for a product catalog — and I have spent years reading how the big ones fail in public. Under the hood there are four jobs that are easy to name and hard to do well: crawl, index, rank, and then the lies we tell about each. If I had to explain the machine to a programmer who will ship a search box on Friday, this is the map I would draw. It is not Google’s paper. It is the parts that still bite when the corpus is your own.

Crawl: you do not have the web. You have a queue and a politeness problem

A crawler is a worker that fetches URLs, extracts links, and puts new URLs on a queue. That sentence hides robots.txt, crawl-delay, canonical tags, infinite calendars, and the site that returns 200 with a soft 404. I have written a crawler that loved a faceted URL scheme and downloaded the same jacket 400 ways. The lie is “we index the web.” Nobody indexes the web. They index a sample they can afford, biased toward what they already think is important.

For a tiny engine I would not crawl the public internet. I would ingest a sitemap, a git repo of markdown, or a product table. If I must crawl my own marketing site, I would start from the sitemap.xml, respect robots, and cap the queue. I would store the raw fetch: status, headers, body hash, fetched_at. Re-crawling without a hash is how you burn budget on unchanged pages.

Rendering is the 2026 argument. A lot of sites are empty HTML plus a JavaScript app. Google runs a renderer. I have used Playwright for a small crawl when the content was client-only. It is slow and it is honest. If your own product docs are static HTML or MDX, skip the browser. If you are building “a search engine” as a startup, the renderer is a product. Pretending fetch() is enough is the first lie.

Diagram-like wall of connected nodes in a dim office

Index: the inverted list is the whole trick, and it is not magic

An inverted index maps a term to the documents that contain it, often with positions so you can do phrases. Lucene, Elasticsearch, OpenSearch, Meilisearch, Typesense, SQLite FTS5 — they are all this idea with different ops stories. I have run Elasticsearch for a catalog and regretted the JVM on a Tuesday. I have used Meilisearch for an indie app and been happier. I have used SQLite FTS when the corpus was a docs site and I did not want a daemon. Elena’s beat is right: a lot of search does not need a cluster.

Tokenization is where languages get political. Lowercasing, stemming, CJK, hyphenated SKUs, code identifiers. If I index invoice_id as two tokens I will miss exact searches. If I do not stem I will miss “invoices.” I pick an analyzer per field: title exact-ish, body stemmed, sku keyword. One analyzer for the whole document is how a search box feels drunk.

The lie is “we found all the pages with that word.” You found the pages you crawled, tokenized, and kept. You dropped robots, noindex, thin content, and the URL you never discovered. Ranking then pretends the remainder is the world.

I also store a forward store — the document I will show in the snippet. People forget the snippet is a second product: you need offsets to highlight. If I cannot highlight, I look like a grep with a UI.

Rank: BM25 is the adult default, and PageRank is not a personality

Lexical ranking in 2026 is still mostly BM25 or a cousin: term frequency, inverse document frequency, field weights, a length norm. I would start there. I would boost title and H1. I would not start with a transformer. I have added embeddings (a small vector field in pgvector or in the search engine) for “this means the same as” after lexical was good. Doing vectors first is how you get a demo that cannot find an exact SKU.

PageRank-shaped signals — incoming links, as a quality prior — still matter on the open web. On a docs site they barely matter. On a store, sales and stock matter more than links. The lie is “we use AI now, so links are dead.” Links are still a cheap prior for spam. AI ranking is a second model that can be wrong in a fluent way. I treat it as a reranker on the top 50, not as the index.

Personalization and freshness are the other knobs. A news query wants time. A medical query wants authority. A product query wants in-stock. One score is a lie. I keep a few rankers and I pick by query class, even if the class is a dumb regex: “looks like a SKU,” “looks like a question,” “looks like a nav.”

Click logs will take over if you let them. They also lock in yesterday’s bias. I will use clicks as a weak signal after I have volume. I will not let them bury a new page that nobody has seen. The cold-start problem is how new documentation dies.

Laptop showing a search box and a list of results

The serving path people skip

Query understanding: spellfix, synonyms, a knowledge of your domain (“void” vs “refund”). I keep a synonym list I can edit without a reindex when I can. I do not hide it in a model I cannot patch on a Friday.

Caching: the same query from the same site should not hit a cold shard every time. I cache the top results with a short TTL. I do not cache personalized results as if they were global. That is how you leak another user’s recent search into a screenshot.

Safety: you will be asked to suppress legal, medical, and spam. The big engines have teams. A tiny engine needs a blocklist and a human. I would rather a blocklist I can explain than a classifier I cannot.

The lies in between, named

  • “We have the whole web.” You have a crawl budget and a politics of discovery.
  • “Keywords are dead.” Lexical match still pays the bills. Vectors help the paraphrase. They do not replace the SKU.
  • “It’s AI now.” A reranker on BM25 is not a new physics. It is a layer. Layers fail independently.
  • “SEO is tricking us.” SEO is a conversation with the crawl and the ranker. Some of it is spam. A lot of it is making the document the crawler can see match the document the user wanted. I will write the programmer SEO piece as a sibling thought; the engine’s job is to reward the match, not the trick.
  • “More index is more better.” A bigger index with worse tokenization is a worse engine. I have deleted documents on purpose and improved NDCG on the queries we cared about.

Spam, SEO, and the adversarial layer

The open web is an adversary. Cloaking, link graphs, generated pages that are fluent and empty. A tiny engine on your own corpus can ignore most of this. The moment you crawl outside, you inherit it. I would add a few cheap heuristics — hidden text, keyword stuffing ratios, a domain age I can store — and I would not pretend they are justice. They are a broom. The broom will hit a good site sometimes. I keep a allowlist for the ones I care about.

Programmers who do SEO are often talking to this layer. Make the document fetchable. Make the title the query. Do not hide the article behind a login the crawler cannot pass. That is not a trick. That is making crawl and index possible. The tricks I would stop doing belong in the SEO essay, not in the engine’s self-image.

If I had to ship a tiny engine this month

Corpus: the product docs and the blog, not the web. Ingest from git. SQLite FTS5 or Meilisearch. Fields: title, headings, body, url. BM25, title boost, a synonym file. A crawler only if something is not in git. No renderer unless I must. A query log so I can see the misses. A weekly twenty-query eval I score by hand. That eval is the product. The engine is the implementation.

If the corpus is a catalog: Typesense or Meilisearch, typo tolerance on, stock as a filter, not as a buried signal. If the corpus is a million legal PDFs: I would pay for OpenSearch and I would budget the ingest. I would not pretend FTS5 is a personality I can scale with hope.

I would not build PageRank. I would not build a web-scale crawler. I would not promise “like Google.” I would promise “finds the page we already wrote.” That promise is still a search engine. It is the only one I have successfully shipped. The rest is a career in spam, geopolitics, and capital. I like the small one. I respect the big one. I do not confuse them in a design review.

More articles for you