Most site search still works the way it did twenty years ago: it looks for the words you typed. Type “something nice to furnish a small apartment” into a classifieds site and you get nothing — no listing contains that sentence. Yet the site has dozens of sofas, shelves and tables that answer it.
Vidifye is a video classifieds marketplace I build and run. This is how its search was rebuilt to find what people mean rather than what they typed, what it takes to run, and what it costs.
Why keyword search fails
Keyword search matches spelling, not meaning. A seller writes “Yamaha MT-07, 2019, low mileage”; a buyer types “cheap motorbike for a beginner”. Same thing, not one word in common. Every search that comes back empty is a buyer who leaves — and on a marketplace, a seller who never hears from them.
What RAG brings to search
RAG — retrieval-augmented generation — is usually discussed for chatbots: first find the right documents, then let an AI answer from them. The first half, the retrieval, is exactly what search needs. Each listing is turned into a vector, a list of numbers that captures what it is about. A search is turned into a vector the same way, and the closest listings come back — even when they share no words with the query.
How a Vidifye search works
- Read the request. An AI model reads the query and pulls out what the person is after: the thing itself, a category, a budget, a condition. “Cheap Yamaha under 5k near Laval” becomes a motorcycle, a ceiling of $5,000 and a place.
- Find by meaning. The cleaned-up request is compared with the vector of every live listing. Each listing is described for this purpose: its title, category, condition, price range, city and description.
- Blend in what matters on a marketplace. Meaning counts most, but not alone: exact words still count, newer listings get a nudge, and nearby ones rank higher. A budget is a hard limit — nothing over $5,000 shows up.
- Take a second look. For longer, descriptive searches, an AI reviews the top ten and reorders them by how well each one actually answers the request. One- or two-word searches such as “honda” skip this step: there is nothing to interpret.
Showing people what was understood
A search that silently reinterprets you is unsettling. So Vidifye shows what it understood: above the results, a line reads “Showing Motorcycles · Under $5,000 · Near Laval”. Each part is a chip that can be cleared in one tap. If the search read the request wrong, the buyer sees it and corrects it — instead of scrolling through results that make no sense and giving up.
That line does more for trust than any amount of ranking accuracy. People forgive a search that misreads them, as long as it tells them how it read them.
Fast enough, cheap enough
The first version took more than four seconds per search — far too slow for someone thumbing through listings on a phone. Three changes brought it down:
- Short searches skip the AI review: “honda” went from 4.6 seconds to 0.85.
- The AI review now runs after the results are sent, and its better order is kept for the next person who searches the same thing.
- The two AI calls at the start run side by side instead of one after the other.
The vectors live in the same MariaDB database as the listings — no separate vector service to pay for, host or secure. Turning all 2,257 existing listings into vectors cost about thirty cents. Repeat searches are answered from a cache, the AI’s reading of a query is kept for thirty days, and each visitor is rate-limited, so the bill follows real buyers, not bots.
Built to fail quietly
Every step can fail: an AI provider has an outage, a quota runs out. When one does, the search carries on with what the previous step produced, and if the AI part is unavailable altogether, the site quietly falls back to classic keyword search. A buyer never sees an error page because a model was busy.
What this means for an SME
If your site has a search box — a catalogue, listings, a knowledge base, support articles — the same approach applies. The pieces are affordable now, the database you already run can probably hold the vectors, and the cost per search is counted in cents.
What makes it work is the engineering around the model: what each item says about itself, how meaning is weighed against price and distance, what happens when the AI is down, and telling people what the machine understood.