AI Search Engine: Onton's Ontology 1 Beat Google and Amazon With 1% of Their Catalogue — No Keywords Required

A San Francisco startup's neurosymbolic model won 52 of 90 intent-heavy shopping queries against Google Shopping (19) and Amazon (16) while indexing roughly 1% of their catalogue. Here is what neurosymbolic search actually means, why the widely-quoted "2.7x more accurate" headline is wrong, and the four things online sellers should do about it.
Key takeaways
- Ontology 1 is a neurosymbolic search model from San Francisco startup Onton, published 29 July 2026 and widely covered on 2–3 August.
- On Subtext-Decor-90 — 90 intent-heavy queries judged by Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5 — Onton won 52 queries outright versus Google Shopping's 19 and Amazon's 16.
- Precision at 10 was 0.630 for Onton, 0.543 for Google Shopping and 0.469 for Amazon: a ~16% relative gain over Google, not the 2.7x many outlets reported.
- Onton indexed roughly 1% of the catalogue its rivals hold, which is the point: structured understanding beat index size on inference-heavy queries.
- The strategic implication for sellers is structured, machine-readable product data — agentic shopping turns noisy product information into financial errors, not just annoyances.
For nearly thirty years, product search has worked the same way: type a couple of keywords, land on a category page, then narrow the results with filters for size, price, brand and material. That model breaks the moment your intent does not map onto a checkbox. There is no filter for "pet-friendly but still looks expensive", and no dropdown for "furniture that might belong in this room".
On 29 July 2026 a small San Francisco company called Onton published Ontology 1, a neurosymbolic AI search engine model built from scratch, and the coverage caught fire this week after independent write-ups landed on 2–3 August. The headline claim is the sort of thing that usually deserves a raised eyebrow: on a benchmark of complex, intent-heavy shopping queries, Ontology 1 beat both Google Shopping and Amazon while having indexed roughly 1% of their catalogue.
This piece does three things. It explains what neurosymbolic search actually means in plain language, it walks through the benchmark carefully — including the number several outlets have reported incorrectly — and it explains what the result implies if you sell products online, build AI agents, or simply want to know where search is heading after keywords.
What Ontology 1 Actually Is
Onton describes Ontology 1 as "a neurosymbolic model that answers complex, intent-heavy searches more accurately than today's leading engines, and improves itself through continuous self-driven learning." Every word in that sentence is doing work, so it is worth unpacking each one.
Neurosymbolic, in plain language
Modern search is dominated by two techniques. Keyword retrieval matches strings and is unbeatable when you know exactly what you want ("Nikon Z6 III body"). Vector retrieval embeds queries and products into a shared numeric space and finds nearest neighbours, which handles paraphrasing well but blurs precise constraints — it is notoriously bad at negation, counting and hard attributes.
A neurosymbolic system adds a third layer: an explicit, structured representation of the world — categories, attributes, relationships, rules — that the neural model reads from and writes to. The neural half interprets messy human language and images. The symbolic half enforces the facts. Ask for "a rug that isn't wool, under $400, that works in a north-facing room" and the symbolic layer can genuinely exclude wool and genuinely apply the price bound, while the neural layer reasons about what "north-facing" implies for colour and warmth.
This is the same architectural instinct we described in our coverage of Astra's machine-checked mathematical proofs: let a model be creative, then let a formal system verify the claim. Different domain, identical philosophy.
Continuous self-driven learning
The second claim is that the model improves itself without a human labelling queue. In practice this means the system generates its own hypotheses about product attributes and relationships, tests them against retrieval outcomes, and promotes the ones that hold. That is what makes the 1%-of-catalogue figure interesting rather than embarrassing: Onton's argument is that a smaller, better-understood index beats a vast, shallowly-modelled one for questions that require inference.
The Benchmark: Subtext-Decor-90
Onton published its methodology alongside the model, which is more than most vendors do. The benchmark is called Subtext-Decor-90, and here is how it works.
- 90 text queries, curated to include rich aesthetic descriptors, negation, cultural references and emotional framing — the cases that defeat keyword and vector retrieval.
- Three engines compared: Onton, Google Shopping and Amazon.
- Three independent multimodal judges: Claude Opus 4.8, Gemini 3.1 Pro and GPT-5.5, each shown screenshots of all three result sets in a single call.
- Metric: precision at 10 (P@10) on the first ten visible result cards, averaged across judges, with 10,000-resample bootstrap confidence intervals across queries.
The results:
- Onton: P@10 of 0.630, 95% CI [0.571, 0.688] — 52 outright query wins.
- Google Shopping: P@10 of 0.543, CI [0.490, 0.596] — 19 outright wins.
- Amazon: P@10 of 0.469, CI [0.417, 0.521] — 16 outright wins.
The Number Being Misreported
Several outlets have run with a headline along the lines of "2.7x more accurate than the world's best e-commerce search engines". That framing does not survive contact with the data, and it is worth being precise about why.
Ontology 1's accuracy advantage on the published metric is 0.630 versus 0.543 — roughly a 16% relative improvement over Google Shopping, and about 34% over Amazon. Respectable, well outside the confidence intervals for Amazon, and clearly meaningful for a product built by a small team. It is not 2.7x.
Where does 2.7x come from? Outright query wins: 52 versus 19 is 2.7x. That is a legitimate statistic about how often one engine produced the best result set for a query, but it is a win rate, not an accuracy multiplier. Conflating the two inflates a good result into an implausible one, and implausible numbers are exactly what makes buyers dismiss a category.
How to read vendor benchmarks: ask what the metric is, who judged, whether confidence intervals overlap, and whether the headline multiple describes the metric or the win rate. Onton's own research post is careful about all four. The coverage often is not.
Three caveats Onton discloses
To its credit, the company flags its own limits. First, precision, not recall — recall is unanswerable without full access to Amazon's and Google's catalogues, so nobody knows what all three engines missed. Second, LLM judges: three frontier models agreeing is a reasonable proxy for human preference on aesthetic queries, but it is a proxy, and models share biases. Third, image and multimodal queries were excluded from the 90 because a one-to-one comparison is impossible — Amazon lacks comparable multimodal search entirely.
That last exclusion cuts both ways. It removes the scenario where Onton claims its widest lead, but it also means the published benchmark understates the differentiator the company is actually selling.
Why "No Keywords" Is the Real Story
The behavioural shift underneath this launch matters more than any single benchmark. Onton cites Bloomreach data indicating that, in 2025, 57% of shoppers used AI to help them shop and 41% now search using natural language rather than keywords. A 2026 Klaviyo study found consumers abandoning keywords in favour of full phrases and questions, with daily AI users routinely issuing queries of eight or more words — sometimes entire paragraphs.
Eight-word queries with negation and emotional framing are a different retrieval problem from two-word queries. Filters cannot express them. Keyword indexes cannot match them. Vector search returns plausible-looking near-misses. This is the gap Ontology 1 targets, and it is why the company built multimodal surfaces — Onton "Canvases", a moodboard-style interface where a picture of your room plus a few words is the query.
Onton's own example is instructive: that query hides judgements about style, scale, palette, what is already in the room and what is missing. No filter touches any of it. This is the same capability curve we tracked in our piece on embodied reasoning models planning in physical spaces — models are getting good at inferring context from images rather than being told it.
The Second Claim: A Trust Layer for AI Shopping Agents
Alongside the accuracy story, Onton positions Ontology 1 as a verification layer: a model built to evaluate whether the product information reaching an AI shopping agent can be trusted. The company argues the industry currently lacks one, and says its internal benchmarks show better product-data accuracy than Google Shopping and Amazon on that axis too.
This is the more strategically interesting claim, because it addresses the failure mode that will define agentic commerce. When a human shops, star ratings, review counts and marketing copy are noisy but survivable — people discount them intuitively. When an autonomous agent shops, that noise becomes input to a purchase decision made without a human reading the page. Wrong dimensions, stale stock, mismatched variants and review manipulation all become financial errors rather than mild annoyances.
Verified product attributes are exactly the kind of capability that gets consumed through a protocol rather than a website. Our guide to MCP servers exposing tools to agents covers the plumbing; a trustworthy product-data endpoint is a natural thing to plug into it.
What This Means If You Sell Online
Ontology 1 itself is not something most businesses will buy this quarter. The transferable lessons are about how discovery is changing, and there are four worth acting on.
1. Structured attributes are becoming your ranking signal
A model that reasons symbolically can only reason over facts it has. Materials, dimensions, care instructions, compatibility, room suitability, weight, country of origin — every attribute you leave in a paragraph of marketing prose is an attribute a neurosymbolic engine cannot use to match you to an intent-heavy query. Machine-readable product data is no longer a feed-management chore; it is discoverability.
2. Long-tail intent is where the traffic is moving
If shoppers are typing eight-word emotional descriptions, your category pages built for two-word keywords are answering a question nobody is asking any more. The practical response is content that names the intent explicitly — use cases, constraints, negations, room and lifestyle framing — rather than another page optimised for a head term. Our AI writing and content hub covers how to do that without producing thin pages.
3. Assume an agent will read your product page
Treat every page as having two audiences: a human and a machine acting for a human. That means schema.org product markup, unambiguous variant data, honest availability and clear returns terms in structured form. The site an agent can parse confidently is the site an agent recommends.
4. Judge new search vendors on your queries, not their benchmark
Subtext-Decor-90 is a decor-heavy benchmark curated by the vendor. That is fine — it is transparent and reproducible — but it is not your catalogue. Build a 50-query set from your own logs, weighted toward the messy, low-converting ones, and score any vendor against it. Our AI for business hub has the evaluation checklist.
The Competitive Picture
It is tempting to read this as a small startup beating two of the largest companies in the world. A more useful reading: Onton is beating them on a slice of the problem they are not optimised for. Amazon's search is tuned for conversion on transactional intent at enormous scale. Google Shopping is tuned for breadth across an index nobody else can assemble. Neither is architected around 90 aesthetic, negation-heavy decor queries.
That does not make the result unimportant. It makes it a classic wedge. The open question is durability: what happens when a hyperscaler applies the same symbolic-plus-neural approach to a catalogue a hundred times larger, and how much of Onton's advantage comes from architecture versus from a carefully curated 1% index. That question will not be answered by another benchmark post — it will be answered by whether shoppers change where they start their search.
Verification Notes
Everything above is drawn from primary and secondary sources you can check yourself: Onton's Ontology 1 research post and its benchmarks write-up with methodology and data, MarkTechPost's coverage, and PPC Land's report on the product-trust angle. For the structured-data work in point three, the canonical reference is schema.org's Product vocabulary. Where a claim is Onton's own and not independently verified — the self-improvement mechanism, the product-trust benchmark — we have said so explicitly.
The Bottom Line
Ontology 1 is a genuinely interesting result reported with an inflated number. The accuracy advantage is about 16% over Google Shopping on the published metric, not 2.7x; the 2.7x figure describes outright query wins. Both are real, and the honest version is still the most credible evidence yet that neurosymbolic retrieval handles intent-heavy natural-language search better than keyword or vector approaches alone.
The strategic signal is bigger than one benchmark. Shoppers have stopped typing keywords, agents are starting to shop on their behalf, and the engines built for a filter-and-category world are the ones with the most to lose. If you sell anything online, the work this quarter is unglamorous and concrete: make your product data structurally true, then test discovery with the messy sentences your customers actually type. Follow the next moves in our generative AI news and AI tool reviews hubs.
Frequently asked questions
What is a neurosymbolic AI search engine?
It combines a neural model, which interprets messy natural language and images, with an explicit symbolic layer of categories, attributes and rules that enforces hard constraints. The neural half handles nuance like "cosy but not rustic"; the symbolic half reliably excludes wool or applies a price ceiling — something vector search does poorly.
Did Ontology 1 really beat Google and Amazon?
On Onton's published benchmark of 90 intent-heavy decor queries, yes: it scored 0.630 precision-at-10 against Google Shopping's 0.543 and Amazon's 0.469, and won 52 queries outright. It is a vendor-curated benchmark in one product category, so treat it as strong evidence for that slice of search rather than a general verdict.
Is Ontology 1 really 2.7x more accurate?
No. The 2.7x figure comes from outright query wins (52 versus Google's 19). The accuracy metric Onton published shows about a 16% relative improvement over Google Shopping and 34% over Amazon. Both statistics are real, but they measure different things and should not be conflated.
How does indexing only 1% of the catalogue still win?
Because the benchmark measures precision in the top ten results, not coverage. A smaller index with rich structured attributes can rank the right ten items for an inference-heavy query more reliably than a vast index with shallow attribute modelling. Recall was not measured, so nobody knows what each engine missed.
What should online sellers do about this?
Four things: expose every product attribute as structured data rather than prose, create content that names long-tail intent instead of head keywords, add complete schema.org Product markup so AI agents can parse your pages confidently, and evaluate any search vendor against 50 queries pulled from your own logs.
Sources & further reading
Every factual claim in this article traces back to the primary sources below. Figures we could not reproduce ourselves are attributed to the vendor in the text.
- Ontology 1 research post — Onton
- benchmarks write-up — Onton
- MarkTechPost's coverage — Marktechpost
- PPC Land's report on the product-trust angle — Ppc
- schema.org's Product vocabulary — Schema
About the author
Way Of Talk Editorial Team — Editorial desk — AI tools, agents and generative AI news
Way Of Talk is written and edited by a small editorial desk that covers new AI tools, agent frameworks and generative AI news. Rather than publishing anonymous content, we publish under a single accountable byline: every article is researched, fact-checked and signed off by the desk, and the desk is reachable at the address below.
Full bio and articles · Editorial policy · editor@timesofai.com
Found this useful? Keep the streak going
We publish a new researched article on the day's trending AI tools topic. Share this piece with a teammate, or jump into another category below.
Browse all articles- #AI search engine
- #no keywords
- #neurosymbolic AI
- #Ontology 1
- #agentic commerce


