Skip to content

What an embedding actually is

An embedding is a list of numbers that puts similar things near each other. That one sentence is enough to build search, recommendations and retrieval — here is how.

4 min read
Contents

The word “embedding” does a lot of damage. It sounds like something you do to a video in a page. What it actually names is the most useful idea in applied machine learning, and you can hold all of it in your head at once.

An embedding is a list of numbers that represents a thing, arranged so that similar things get similar lists.

That is it. Everything that follows is consequences.

Start with something you can picture

Forget text for a second. Imagine describing cities with two numbers: latitude and longitude.

const cities = {
  dubai:   [25.2, 55.3],
  sharjah: [25.3, 55.4],
  karachi: [24.9, 67.0],
  lisbon:  [38.7, -9.1],
};

Now “which city is most like Dubai?” becomes arithmetic. Subtract the pairs, measure the distance, take the smallest. Sharjah wins, Karachi second, Lisbon nowhere. You did not teach the computer anything about cities. You chose a representation where closeness in numbers means closeness in reality, and similarity fell out for free.

An embedding model does the same job for things that have no natural coordinates. Give it a sentence, it hands back a list of numbers — typically 384, 768 or 1536 of them instead of two. You cannot picture 768 dimensions and you do not need to. The rule is identical: sentences about similar things land near each other.

What “near” means in code

Two vectors, one number out. The standard measure is cosine similarity, which asks whether two vectors point the same direction and ignores how long they are:

function cosineSimilarity(a, b) {
  let dot = 0;
  let magA = 0;
  let magB = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    magA += a[i] * a[i];
    magB += b[i] * b[i];
  }
  return dot / (Math.sqrt(magA) * Math.sqrt(magB));
}

Twelve lines. It returns roughly 1 for “these mean the same thing”, roughly 0 for “unrelated”, and negative for opposites, though in practice you rarely see a genuinely negative score with modern text models.

That function plus an embedding model is a complete semantic search engine. Everything else is performance.

A user types “my card got declined”. Your knowledge base article is titled “Troubleshooting failed payments”. A LIKE '%card declined%' query finds nothing. Full-text search finds nothing, because the two strings share no useful words.

Embed both and they land close together, because the model learned from enough text to know that declined cards and failed payments are the same neighbourhood. That is the entire pitch for vector search, and once you have seen it work on a real support site it is hard to go back.

The three things that actually bite you

I have shipped this a few times now, and the same three problems show up every time.

1. Chunking decides your quality, not the model

You cannot embed a 4,000-word article as one vector and get anything useful. The meaning smears. Split it, and now every decision matters: split by paragraph and you lose the context a paragraph sits in; split by heading and one section might be 3,000 words long.

What has worked for me: split on headings, then split anything still over ~500 words on paragraph boundaries, and prepend the document title and heading path to every chunk before embedding it.

const chunkText = `${doc.title} › ${section.heading}\n\n${section.body}`;

That prefix costs you a few tokens and buys back most of the context you lost. It is the highest-leverage change in the whole pipeline, and it is three lines.

2. You must use the same model for queries and documents

Two embedding models produce numbers in completely unrelated coordinate systems. A vector from model A compared against a vector from model B returns noise that looks like a plausible score, which is the worst possible failure — it does not crash, it just quietly ranks badly.

Store the model name and version alongside every vector. When you change models, you re-embed everything. Budget for that from day one, because you will change models.

3. Storage is a real decision

For a few thousand items, a column in Postgres and a loop is genuinely fine. I have shipped that. At tens of thousands you want an index:

-- pgvector: an index that trades a little recall for a lot of speed
CREATE INDEX ON documents
  USING hnsw (embedding vector_cosine_ops);

Postgres with pgvector covers almost every project a web developer will build. Reach for a dedicated vector database when you have a specific reason, not because a blog post said to.

Things you can build with this today

Embeddings are not only for chatbots. Some of the most useful features I have shipped with them have nothing to do with generating text:

  • Related posts that actually relate, instead of sharing a tag.
  • Duplicate detection on support tickets, before a human reads the second one.
  • Routing, where an incoming message is compared against a handful of example messages per department and sent to the nearest.
  • Clustering — embed a month of feedback, group the vectors, and read one example from each group instead of all 800.

That last one takes about forty lines and has changed more product decisions at my clients than any dashboard I have ever built.

The honest limitation

Embeddings capture topical similarity, not truth and not logic. “The deployment succeeded” and “the deployment failed” are extremely close in embedding space, because they are about the same thing. If your feature depends on telling those two apart, similarity is the wrong tool and you need a model that actually reads.

Know what the number means, and it will serve you for years. Treat it as understanding, and it will embarrass you in front of a client.

Share