Cloudflare AI Gateway adds managed web search for retrieval-augmented applications
Cloudflare AI Gateway now offers managed web search through Ceramic.ai, Exa, and Linkup. This article breaks down when the managed API wins and when a custom retrieval pipeline still pays off.
Cloudflare AI Gateway now offers a native web search API through partnerships with Ceramic.ai, Exa and Linkup, letting developers inject real-time web context into model calls without building custom middleware. For teams running retrieval-augmented generation pipelines, this shifts the build-versus-buy decision: managed search removes the need to operate crawlers, indexers and ranking logic, but it also means accepting the providers' source coverage, latency profiles and pricing.
Managed search replaces the middleware layer
Most retrieval pipelines today stitch together a search API, a crawler or browser automation layer, content extraction, chunking, embedding and a vector store before the model ever sees context. Cloudflare's Web Search API collapses that stack into a single call that returns structured snippets with source URLs. The request flows through AI Gateway, so logs, billing, access controls and rate limits are handled in the same control plane used for model inference [^1].
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/ \
--request POST \
--header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"query": "What are some fun things to do in Salt Lake City as fall approaches?",
"provider": "ceramic",
"limit": 5,
"options": { "gateway": { "id": "default" } }
}'
On Workers the same call is a one-line binding:
const response = await env.AI.websearch({
gatewayId: "default",
query: "What are some fun things to do in Salt Lake City as fall approaches?",
provider: "exa",
limit: 5,
})
const results = await response.json()
The gateway also supports Bring-Your-Own-Key for teams that already have contracts with the search providers, and Cloudflare says it adds no markup on top of the providers' list pricing [^1].
When the managed approach wins
A managed search API makes sense when:
- Freshness matters more than depth. The providers return live snippets, which solves the knowledge-cutoff problem for recent events, changing APIs or fast-moving news. Cloudflare's launch post highlights retrieving the latest documentation released during its Birthday Week as a canonical example [^1].
- Team capacity is limited. Operating a crawler fleet that respects
robots.txt, handles JavaScript-rendered pages, deduplicates results and stays within rate limits is a dedicated engineering effort. Cloudflare requires its partners to meet its "Verified bots" standard — identified crawlers,robots.txtcompliance and source attribution in every response [^1]. - Observability and cost control are already centralised on AI Gateway. Search calls appear in the same logs as model calls, draw from the same credit balance and inherit the same access policies. Zero Data Retention flags are surfaced per provider so compliance teams can verify data handling without extra tooling [^1].
When a custom pipeline still pays off
A self-built pipeline remains the better choice when:
- Domain coverage is narrow and deep. Legal filings, regulatory gazettes, specialised technical forums or paywalled industry sources often sit outside general web indexes. If the application depends on sources the three providers do not crawl, a custom crawler with authenticated access is unavoidable.
- Retrieval quality must be tuned per query type. Hybrid search (keyword + vector), reranking with a cross-encoder, query rewriting and recursive retrieval all require control over the index and ranking pipeline that a black-box API does not expose.
- Latency budgets are tight and predictable. A managed API adds a network hop and the provider's internal latency. Custom pipelines can cache aggressively, pre-fetch known hot documents and serve from a local vector store with sub-10 ms p99.
- Cost scales non-linearly with volume. At high query volumes the per-call price of a managed API can exceed the amortised cost of a self-hosted index, especially when the same documents are retrieved repeatedly.
Latency and cost trade-offs in practice
Cloudflare does not publish latency SLAs or per-provider benchmarks in the launch material. The providers differ in architecture:
- Ceramic.ai positions itself as a research-grade search engine with a focus on citation quality.
- Exa (formerly Metaphor) uses a neural index optimised for semantic similarity over keyword matching.
- Linkup emphasises structured data extraction and API-friendly responses.
Because the API is routed through AI Gateway, each search call incurs the gateway's overhead plus the provider's latency. Teams should run a small A/B test: route a percentage of production queries to each provider, measure end-to-end latency (gateway + provider + model), and compare against a baseline custom pipeline on the same queries. The gateway logs make this straightforward — every search call appears alongside the model call that consumed its results [^1].
Pricing is pass-through at the providers' list rates. As a rough guide, public pricing for comparable web search APIs ranges from $3–$10 per 1,000 queries depending on depth and freshness guarantees. Multiply by your expected daily query volume and compare against the engineering cost of maintaining a custom pipeline (crawler infrastructure, index storage, embedding compute, on-call rotation).
Evaluating source coverage for your domain
The launch post does not publish the crawl corpora of Ceramic, Exa or Linkup. To evaluate coverage:
- Sample representative queries from your production logs — product questions, troubleshooting searches, regulatory lookups, competitive intelligence.
- Run each query against all three providers via the REST API or Workers binding and inspect the returned URLs.
- Check domain overlap with your known authoritative sources. If a provider consistently returns the same 10–20 domains your custom crawler would target, coverage is likely sufficient.
- Verify recency by querying for events or releases from the past 24–48 hours and confirming the snippets reflect the latest version.
- Test edge cases: paywalled content, JavaScript-heavy sites, PDF-heavy repositories, non-English sources.
If gaps appear, the managed API can still handle the long tail of general web queries while a custom pipeline covers the high-value verticals.
Server tools are coming — but you can orchestrate today
Cloudflare plans to ship native Server Tools inside AI Gateway, with web search as one of the first built-in tools. Until then, the launch post shows a complete Workers pattern: the model calls a web_search function, the Worker executes the search binding, and the results are fed back into a second model call [^1]. This pattern works today and mirrors the OpenAI function-calling convention, so existing agent frameworks can adopt it with minimal changes.
What to do next
- Enable AI Gateway if you have not already — it is the control plane for the search API.
- Run a coverage audit against your top 50 production queries using all three providers.
- Benchmark latency and cost at your expected query volume, including the gateway overhead.
- Decide the split: managed search for general freshness, custom pipeline for domain-critical sources.
- Instrument the gateway logs to monitor search success rates, provider latency and credit consumption from day one.
A & A Labs helps teams put large language models into production over their own data with retrieval evaluation, cost and latency controls, and human-review checkpoints so answers stay accurate, source-attributed and affordable to run.
Frequently asked questions
Which search providers does Cloudflare AI Gateway support for web search?
Cloudflare AI Gateway supports three search providers: Ceramic.ai, Exa (formerly Metaphor), and Linkup. Each has different architectural focuses — Ceramic on citation quality, Exa on semantic similarity via neural index, and Linkup on structured data extraction.
How much does Cloudflare's managed web search API cost?
Cloudflare passes through the providers' list pricing with no markup. Public pricing for comparable web search APIs ranges from $3–$10 per 1,000 queries depending on depth and freshness guarantees. Teams with existing contracts can use Bring-Your-Own-Key.
When should I use Cloudflare's managed web search instead of a custom RAG pipeline?
Use managed search when freshness matters more than depth (live snippets solve knowledge cutoff), team capacity is limited (no crawler fleet to operate), or observability and cost control are already centralised on AI Gateway (shared logs, billing, access controls, Zero Data Retention flags).
When is a custom retrieval pipeline still the better choice?
Custom pipelines win when domain coverage is narrow and deep (legal filings, regulatory gazettes, paywalled sources outside general indexes), retrieval quality must be tuned per query type (hybrid search, reranking, query rewriting), latency budgets are tight and predictable (sub-10ms p99 from local vector store), or cost scales non-linearly with volume (per-call price exceeds amortised self-hosted index cost).
How do I evaluate which provider covers my domain best?
Sample representative queries from production logs, run each against all three providers via REST API or Workers binding, inspect returned URLs for domain overlap with authoritative sources, verify recency with 24–48 hour events, and test edge cases like paywalled content, JavaScript-heavy sites, PDF repositories, and non-English sources.