They all return the page. One tells you the rules.
A map, not a scoreboard. Every figure below was read from that company's own public pages and carries its source, and further down we list in plain words what we are not claiming.
Our index size in the tables below is read from our public counter at api.lyrenth.com/v1/stats every time this page is built, so it is never more than a minute old. Everything else on this page is dated instead.
Retrieval APIs.
Sold per request to agents and developers, and several run no web-scale crawl at all: they fetch on demand. Two of the four own an index, one owns none and says so plainly, and one does not state where its general web results come from. The independent indexes are a different business, and have their own table below.
| Capability | Lyrenth | Exa | Firecrawl | Tavily | Jina |
|---|---|---|---|---|---|
| Owns its own crawler and index | Yes, and gated on nobodyour own crawler fleet, our own index, our own ranking. No part of our crawling, indexing, or ranking depends on another engine's results or permission: we crawl what a site's own robots.txt allows, not what a third party allows. The counter above is live, not a number we typed in | Yesown index and crawler (ExaSearchBot) | Not publishedthey name specialist indexes (research papers, developer docs) but never state where general web search results come from, so we assert neither own index nor reseller | Yesown crawler and index, stated | No indexstated plainly: the Reader is not a consumer search engine and does not index or rank the web on your behalf. Live fetch per request, 5 minute cache |
| Published index size, in their own qualifier | 2,998,608,191documents indexed right now, not a claim we typed in and not a figure with a date on it. Growing by over 50 million documents a day, measured over the past week. At that rate the index passes 9 billion by the end of 2026, which is where the nearest independent index stands today | 80 billiontheir qualifier: documents served in their vector DB. Separately, 1.4 trillion URLs tracked. No date given | Not publishedno index-size claim on their pages | Not publishedno index-size claim. Their billions of pages crawled is throughput, not corpus, so it is not reported here as an index size | No indexnothing to publish a size for |
| Search price | Built, not livethe search API is written and tested behind a gate. When it opens, searching is free on every plan within a per-tier fair-use allowance and the read remains the only paid unit. Every other column here bills the query whether or not you use the answer | $7 / 1,000search requests. Answer endpoint $5 / 1,000 | Subscription onlyno pay-per-use. Hobby $19/mo for 5,000 credits, Standard $99/mo for 100,000, Growth $399/mo for 500,000, Scale $749/mo for 1M. Search costs 2 credits per 10 results | $0.008 / creditpay as you go. Their paid tier is set on a slider rather than fixed steps, so there is no tier list to quote (their pricing page, re-read 1 September 2026) | 10,000 tokensfixed minimum per search request, billed from the token balance |
| Read or extract price | $0.40 to $0.95 / 1,000reads, by plan, drawn from the account's credits. Starter $19 a month for 20,000 reads is $0.95 per 1,000; Pro $79 for 200,000 is $0.40. The free tier is 2,000 reads a month with no card | $1 / 1,000contents (pages) | Plan creditsno pay-per-use rate published for a read | Not publishedno separate extract rate published | $50 / 1B tokensand $500 for 11B. Usable with no key at 20 requests per minute |
| Free tier | Yes2,000 reads a month, no card, and 25 an hour with no account at all | Yes$20 of credits on signup plus $10 monthly, both as published. Their page does not reconcile the two, so neither do we | Yes1,000 credits, no card required | Yes1,000 credits monthly, no card required | Yes, non-commercial10M tokens free, but under CC-BY-NC, so the free tier is not usable commercially |
| Crawler a site owner can identify and block | Yes, all fouruser agent AIWebIndex/2.0, published IP ranges at /bot/ip-ranges.json, forward-confirmed rDNS under lyrenth.com, and Web Bot Auth signatures (RFC 9421) | Signed, no IP listuser agent plus Web Bot Auth signatures (RFC 9421, Ed25519), the same mechanism we use. No IP ranges, by design: the signature holds regardless of IP | Not publishedno crawler identity on their own pages. Third-party trackers name a string, and open GitHub issues ask them to document one, so we do not print a name they have not published | No, by policythey state they do not advertise a differentiated user agent, to avoid discrimination by sites that allow only Google, and will not crawl what googlebot cannot | Nono crawler identity, and callers can override the user agent, which cuts against identifiability. They do state they do not bypass anti-bot defences |
| Pays publishers | Built, not livethe design is self-serve: a verified site owner sets their own per-read price, from $0.001 to $0.05, and keeps 73% of every paid read of their pages, with no partner negotiation, no traffic minimum and no invitation. It is built and not open yet, so no site owner is being paid through it today | NoExa Connect pays commercial data vendors, not crawled sites | Nostated intent in their Series A post, no live program | Nono program found | Nono program found |
| The clean document is ready before you ask | Already built and storedthe clean document is produced when we crawl the page and stored whole: markdown, headings, links, media, quality signals, and the rights state. A read is a lookup of something that already exists, not a fetch and a parse run while you wait, and one crawl serves every caller | Not publishedthey run their own index and crawler and sell page contents at $1 per 1,000, but their own pages do not state whether those contents come from storage or are fetched when you ask | Not publishedno index-size claim, and no statement of where general web results come from, so their own pages do not establish what exists before a request | Not publishedown crawler and index stated, but no read or extract product is published, so what a caller receives is not stated either | Fetched when you askstated plainly: no index, a live fetch per request, with a 5 minute cache. The document is produced while you wait rather than read from an index |
| Only one yes in this rowPer-document licence and rights in API responses | Yesevery AIDocument carries the document's licence state: none, attribution, paid_required, or noai, and search results will carry the same field when search opens | Nocontent returned with no rights field | Nocontent returned with no rights field | Nocontent returned with no rights field | Nocontent returned with no rights field |
| Funding posture | Bootstrappedno venture funding | Venture funded$250M Series C at $2.2B (third-party report) | Venture funded$14.5M Series A, stated on their own blog | Being acquiredacquisition by Nebius, announced 10 February 2026 on their own blog and the Nebius newsroom, so no longer an independent startup | Venture funded$30M Series A from 2021 (third-party report, likely stale) |
Independent indexes.
Own crawler, own index, own ranking, all three built in house. It is a short list, and this is the group we belong to. We add over 50 million documents a day, measured over the past week, so the index passes 9 billion by the end of 2026, which is where the nearest name in this table stands today. Of the four columns here, only ours sends the rights with the content and pays site owners self-serve.
| Capability | Lyrenth | Brave | Mojeek | Ahrefs / Yep |
|---|---|---|---|---|
| Owns its own crawler and index | Yes, and gated on nobodyour own crawler fleet, our own index, our own ranking. No part of our crawling, indexing, or ranking depends on another engine's results or permission: we crawl what a site's own robots.txt allows, not what a third party allows. The counter above is live, not a number we typed in | Yes, with two qualifiersthey state independence from Google and Bing, and also publish that the index is fed partly by opt-in browser telemetry (the Web Discovery Project) and that their crawler will not crawl a page googlebot cannot crawl | Yestheir wording: results 100% independent, own crawler, own index, own ranking, and no part of indexing or ranking based on results from other engines | YesAhrefsBot powers both Ahrefs and Yep, with a stated 8 billion page daily crawl |
| Published index size, in their own qualifier | 2,998,608,191documents indexed right now, not a claim we typed in and not a figure with a date on it. Growing by over 50 million documents a day, measured over the past week. At that rate the index passes 9 billion by the end of 2026, which is where the nearest independent index stands today | Over 30 billion pagestheir wording, no date given, plus over 100 million page updates every day | Over 9 billion pages (2025)their own timeline dates passing 9 billion to 2025 and 8 billion to 2024. No 2026 update exists, so this may understate them today | 100 billionthe Yep search index, corroborated on two of their own pages, no date stated. Their 493.9 billion figure is the backlink crawl, a different measure, and is not used here |
| Search price | Built, not livethe search API is written and tested behind a gate. When it opens, searching is free on every plan within a per-tier fair-use allowance and the read remains the only paid unit. Every other column here bills the query whether or not you use the answer | $5 / 1,000search requests at 50 queries per second. Answers $4 / 1,000 plus $5 per million tokens at 2 queries per second. Storing results requires a plan that grants storage rights | £2 / 1,000 queriesStartup tier, excluding VAT, at 5 queries per second and 100,000 a day with up to 10 results. Business £3 / 1,000 at 10 per second and 400,000 a day with up to 40 results. Enterprise is custom | $4 / 1,000Yep Search API, Balanced tier. Advanced $8 / 1,000. First 20 results included, each additional result $1 / 1,000, charged only for results returned |
| Read or extract price | $0.40 to $0.95 / 1,000reads, by plan, drawn from the account's credits. Starter $19 a month for 20,000 reads is $0.95 per 1,000; Pro $79 for 200,000 is $0.40. The free tier is 2,000 reads a month with no card | No read productsearch results only. Their API FAQ states it grants no rights to third-party content | No read productweb search API only | No read productthe Ahrefs API v3 is SEO data, not web search, and its dollar pricing is not published in a form we could verify, so no number is shown for it |
| Free tier | Yes2,000 reads a month, no card, and 25 an hour with no account at all | Credits, card required$5 of free credits monthly on both plans. There is no zero-cost plan and a card is required | Free trialquery limit not specified | Yes, 1,000 requestsno credit card required |
| Crawler a site owner can identify and block | Yes, all fouruser agent AIWebIndex/2.0, published IP ranges at /bot/ip-ranges.json, forward-confirmed rDNS under lyrenth.com, and Web Bot Auth signatures (RFC 9421) | No, by policytheir crawler help page says it does not advertise a differentiated user agent, to avoid discrimination by sites that allow only Google. No token, no IP ranges, no rDNS method, so a site owner cannot identify or selectively block it | YesMojeekBot token, a published IP list, forward-confirmed rDNS under mojeek.com with a worked example, obeys robots.txt and the noindex, nocache, and nofollow tags, and caps at one page per site per second | Yespublished AhrefsBot and AhrefsSiteAudit user agents, a robots.txt token, crawl-delay support, published IP range endpoints, an rDNS suffix rule of ahrefs.com or ahrefs.net, and a Cloudflare verified bot listing |
| Pays publishers | Built, not livethe design is self-serve: a verified site owner sets their own per-read price, from $0.001 to $0.05, and keeps 73% of every paid read of their pages, with no partner negotiation, no traffic minimum and no invitation. It is built and not open yet, so no site owner is being paid through it today | Nono index-licensing or publisher-payment program found. Brave Rewards and Brave Creators are browser products, not search | Nono program found | Announced 2022, none published todayYep announced an advertising revenue share with creators at launch in 2022 (third-party report). Their own current timeline describes it in the past tense and describes a 2025 pivot to the Search API, and what site owners are offered today is stated as reach, not money. We do not claim they pay publishers today, and we do not claim they cancelled it |
| The clean document is ready before you ask | Already built and storedthe clean document is produced when we crawl the page and stored whole: markdown, headings, links, media, quality signals, and the rights state. A read is a lookup of something that already exists, not a fetch and a parse run while you wait, and one crawl serves every caller | Search results onlyno read product, so the API returns results and fetching and parsing each page is left to the caller | Search results onlyweb search API only, so fetching and parsing each page is left to the caller | Search results onlythe Yep Search API returns results, and the Ahrefs API v3 is SEO data rather than page content, so fetching and parsing each page is left to the caller |
| Only one yes in this rowPer-document licence and rights in API responses | Yesevery AIDocument carries the document's licence state: none, attribution, paid_required, or noai, and search results will carry the same field when search opens | Nosearch results carry no per-document rights field | Nono per-document rights field. Worth noting the opposite of Brave on a related point: storage rights and AI usage are included on every tier | Nosearch results carry no per-document rights field |
| Funding posture | Bootstrappedno venture funding | Venture and token-sale fundedamounts not published: no primary source exists | Angel fundedinvestors are private individuals, none institutional, and they state they have not and will not take venture capital (stated 2020) | Bootstrappedin their own words on their about page. They are a genuinely self-funded index, so we are not the only one |
Rights travel with the document.
Every search result and every AIDocument carries the document's machine-readable licence state: none, attribution, paid_required, or noai, with the source it was learned from and a terms URL when one exists. An agent can decide whether it is allowed to use what it just fetched without a lawyer in the loop.
Checked against every provider here on 14 August 2026, and not redone since: none of them surfaces a per-document rights field in API responses. Content comes back, rights do not. That is the one row where this column stands alone, and it is what the whole page is built around.
- Not the biggest. Ahrefs publishes 100 billion for the Yep index and Brave publishes over 30 billion pages. We publish a live counter and it is smaller.
- Not the only bootstrapped one. Ahrefs is self-funded in their own words, and Mojeek is angel funded and states it has not and will not take venture capital.
- Not the first to pay the web back, and ours is not open yet. Others compensate content owners today, through curated partner sets and platform-computed models. Ours is built and switched off, and its difference is shape: any verified owner sets their own per-read price and takes 73%, self-serve.
- Not the only verified crawler. Exa signs its requests with the same Web Bot Auth mechanism (RFC 9421), and Mojeek and Ahrefs publish user agents and IP lists. Our claim is completeness: user agent, published IP ranges, forward-confirmed rDNS, and signatures, together.
One is live. Three are built and switched off.
Stated exactly, because the opposite error is the expensive one: the index and the read serve customers today. Search, paying the web back, and the Verified Read are written and tested behind gates, and nobody is using them yet.
Live today: the index and the read
Search the index and the same result resolves into a clean AIDocument in one product. The document count on this page is the live counter, not a figure typed into a slide.
Built, not live: search
Written and tested behind a gate, not switched on. One query will return ranked, rights-aware sources, and any result resolves into a clean AIDocument in the same product. The API speaks HTTP QUERY (RFC 10008) alongside POST, so identical searches can cache at the edge.
Built, not live: paying the web back
Built and not open yet, so nobody is being paid through it today. The design is self-serve: a verified site owner sets their own per-read price and receives 73% of every paid read, with no partner negotiation and no invitation.
Built, not live: the Verified Read
Built and dark, so no advertiser is buying and no sponsored result is being served. The design prices on the Verified Read, the moment an agent actually reads the document, rather than on an impression nobody saw. Sponsored results would arrive in their own labelled array and never change the order of unpaid results.
The read is the only unit we charge for, and it is the one that is live. Current read rates and plan limits live on the pricing page, so they cannot go stale here.
Where every competitor figure came from.
Every page listed below was opened on 14 August 2026, and the note on each row says which figure came from which page. On 1 September 2026 we re-read them: Brave, Mojeek and Jina were unchanged, Firecrawl's prices were unchanged behind a page that now displays them billed annually, and Tavily had replaced its fixed tiers with a slider, so that cell was corrected. Yep is the one we could not re-read, because their site refuses our crawler, so their figures remain as they were on 14 August.
Competitor pricing and index claims change on their schedule, not ours. Everything above was read from each company's own public pages, researched on 14 August 2026 and re-read on 1 September 2026, quoted in their own qualifier and shown with the source it came from. One we could not re-read: Yep refuses our crawler, so their figures remain the August reading and are marked as such above. Where a company does not publish a figure the cell says Not published rather than an estimate, and where a figure was reported by someone other than the company it is marked as a third-party report. Third-party trackers and aggregators were not used as a source for anything stated as fact.
Where a provider is listed by domain rather than by a specific page, the figures came from that company's own pricing and documentation pages, read on the same date. If we have a figure wrong or it has moved since that date, tell us on the contact page and it gets corrected.