word2vec isn’t just for words (daily search tip)


In my previous tip I introduced word2vec. I discussed it in terms of language: this word, mary, shared context with this other word, lamb, so their embeddings move closer.

Why constrain ourselves to language?

We could pretend that “Doug likes Star Wars” is the same kind of co-occurence. We can make a table of users to the movies they like:

Anchor Positive movie Negative movie
doug star wars king kong
doug star trek cinderella
tom star wars citizen kane
tom battlestar galactica the aviator

Think about what we have:

  • Doug and Tom’s embeddings grow closer through star wars. A word2vec training here shrinks the distance from Doug ←→Star Wars and Tom ←→ Star Wars, making Doug a more similar user to Tom.
  • In the same way, battlestar galactica moves closer to star trek through doug + tom

Thus now, we have a movie recommender system, through the same technology behind word2vec.

We could use this for quite a lot of domains:

  • Queries and documents
  • Images and captions

And so on!

-Doug

PS - 5 days left to signup for Cheat at Search with Agents!

Events · Consulting · Training (use code search-tips)

You're subscribed to Doug Turnbull's daily search tips where I share tips, blog articles, events, and more. You can always manage your profile:

Doug Turnbull

I share search tips, blog articles, and free events I'm hosting about the search+retreval industry, vector databases, information retrieval and more.

Read more from Doug Turnbull

My work is split between mature search teams and new AI teams. Search teams are often farther along and manage a mature, traditional search product. AI teams, however, often don't know what they don't know yet. They've just been cobbled together, have built a few agent demos, and are often in the process of discovering the three big mistakes I blog here: Evals+measurement need to dominate a lot of your product thinking Retrieval isn't "one thing" (ie classic RAG) - its extremely custom to...

I'm writing about the weak spots in vector databases. Where you should prod and poke when selecting a vendor. Today: Updating Vector Databases If you think about the old vector search regime, it involved Indexing everything up front Never updating the index Search-only We overindexed on this paradigm, creating data structures focused on good search performance that couldn't tolerate updates. I wrote about how sensitive graph-based vector DBs in particular are to updates, picking on Lucene...

Hey all, I wrote a new article about a technique that has come up over and over in my work, especially in my Cheat at Search training, for doing effective query understanding into a large vocabulary. Instead of asking an LLM to classify into a vocabulary. Ask it to hallucinate fake entities, then resolve those to real ones client side. Save yourself a lot of tokens and use cheaper models. https://softwaredoug.com/blog/2026/08/10/hypothetical-classifications -Doug PS - A reminder that Vectors...