---
title: "How DoorDash moved from vendor-first to open-weight models in its GenAI platform"
canonical: https://aa-labs.co/blog/how-doordash-moved-from-vendor-first-to-open-weight-models-in-its-genai-platform
author: A & A Labs
published: 2026-10-07
updated: 2026-10-07
category: Engineering
publisher: A & A Labs
reading_minutes: 5
---

# How DoorDash moved from vendor-first to open-weight models in its GenAI platform

> DoorDash's GenAI Platform team moved from vendor APIs to self-hosted open-weight models to gain control over model versions, achieve predictable costs through reserved GPU capacity, keep data inside their VPC for compliance, and reuse a single evaluation harness across model changes. They built an LLM Gateway and Agent Gateway that abstract model endpoints, enabling gradual migration without rewri

## Key takeaways

- DoorDash shifted from vendor APIs to self-hosted open-weight models to control model versions, costs, and data residency.
- An LLM Gateway and Agent Gateway abstract model endpoints, enabling gradual migration without application rewrites.
- A reusable evaluation harness with golden datasets and promotion gates was the key enabler for low-risk model upgrades.
- Reserved GPU capacity turns variable per-token spend into fixed OpEx, amortised across 5,000+ internal users.
- A five-criterion scorecard helps teams decide when to move from vendor APIs to self-hosted inference.

DoorDash's GenAI Platform team shifted from vendor APIs to self-hosted open-weight models to gain control over cost, data residency and evaluation. The architectural bets they made — model ownership, predictable pricing, compliant data handling and a reusable evaluation framework — give engineering managers a repeatable decision framework for their own internal AI platforms.

## The customer changed: from ML engineers to every engineer

When the GenAI Platform team formed under the ML Platform organisation in April 2023, they assumed their users would be machine learning engineers. A conversation quickly corrected that. A product engineer asked what a notebook was, and the team realised the audience was *all* engineers — and, as adoption grew, 40 % non-engineers from legal, sales, operations and strategy [source](https://www.infoq.com/presentations/doordash-genai-platform-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global).

That insight drove two early decisions. First, the platform would be API-first and SDK-first, not notebook- or GPU-first. Second, the team would optimise for business impact — automation that reduces cost, and recommendations or personalisation that grow revenue — rather than chase chatbots or coding agents. Their unique value became helping product teams navigate the accuracy–latency–cost triangle on every use case.

By QCon AI 2025 the platform had over 5,000 internal users, with roughly 45 new users onboarding every day [source](https://www.infoq.com/presentations/doordash-genai-platform-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global).

## Adoption created accountability: the vendor-first starting point

The initial vendor-first approach made sense in April 2023. Signing the OpenAI contract was the fastest way to unblock product teams, and the spend "seems minor now" compared with later scale [source](https://www.infoq.com/presentations/doordash-genai-platform-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global). But rapid adoption turned convenience into liability:

- **Cost predictability** — per-token pricing scales non-linearly with usage; finance teams need committed, forecastable spend.
- **Data residency and compliance** — vendor APIs send prompts and completions outside the organisation's trust boundary, creating legal and regulatory friction for PII, payment data and proprietary business logic.
- **Model ownership** — vendor deprecations, rate-limit changes and opaque model updates break downstream prompts and evaluations without notice.
- **Evaluation portability** — prompt templates, few-shot examples and guardrails built for one vendor's model family do not transfer cleanly to another.

The team recognised that accountability for accuracy, latency and cost meant owning the serving stack, not just the prompt layer.

## The pivot: self-hosted open-weight models

DoorDash's transition followed a staged migration rather than a big-bang cutover.

### 1. Model ownership as a strategic lever
Self-hosting lets the platform team pin model versions, control quantisation levels and schedule upgrades behind feature flags. When a new open-weight release (for example, a Llama 3.x variant) shows better instruction-following on DoorDash's internal eval sets, the team can A/B test it in the gateway before promoting it to default. Vendor APIs offer no equivalent control.

### 2. Cost predictability through reserved capacity
GPU clusters — whether on-prem or reserved cloud instances — turn variable per-token spend into a fixed OpEx line item. The platform team can amortise capacity across 5,000+ users and multiple use cases, smoothing peaks that would otherwise trigger vendor rate limits or surprise invoices.

### 3. Data residency by design
Self-hosted inference keeps prompts, retrieval context and completions inside DoorDash's VPC. This satisfies data-processing agreements with merchants, dashers and regulators without needing vendor DPAs or data-processing addenda that often lag product launches.

### 4. Evaluation infrastructure that travels with the model
The team built a reusable evaluation harness: golden datasets per use case, automated metric suites (accuracy, hallucination rate, latency percentiles, cost per 1k tokens) and a promotion gate that blocks model upgrades unless they meet thresholds. Because the harness sits above the model endpoint, it works identically for vendor APIs during the transition and for self-hosted models afterward. This was the single biggest enabler of a low-risk migration.

## Gateway and agent layer: the abstraction that made migration possible

The GenAI Platform exposes an **LLM Gateway** and an **Agent Gateway** [source](https://www.infoq.com/presentations/doordash-genai-platform-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global). Product teams call a stable interface (`/v1/chat/completions`–compatible for LLMs, a task-oriented RPC for agents) while the gateway handles:

- Request routing (vendor vs. self-hosted, model version, quantisation tier)
- Authentication, rate limiting and quota enforcement per team
- Observability: structured logs, distributed traces, cost attribution tags
- Fallback chains (e.g., self-hosted primary → vendor backup on capacity exhaustion)
- Prompt-template versioning and feature-flagged guardrails

Because the gateway absorbs model-specific quirks (tokeniser differences, stop-sequence behaviour, context-window limits), product teams migrated use cases one at a time without rewriting application code.

## Decision framework you can apply today

| Criterion | Vendor API | Self-hosted open-weight | DoorDash's threshold for switching |
|-----------|------------|-------------------------|-------------------------------------|
| **Model ownership** | Vendor controls version, deprecation, quantisation | Team controls all three | Need to pin versions for eval stability |
| **Cost predictability** | Per-token, variable | Fixed capacity, amortised | Monthly spend > reserved-cluster breakeven |
| **Data residency** | Data leaves trust boundary | Data stays in VPC | Regulatory or contractual requirement |
| **Evaluation portability** | Prompt/guardrail tied to vendor model family | Harness works across any OpenAI-compatible endpoint | Reusable eval harness already built |
| **Operational maturity** | Zero ops burden | Requires GPU fleet, autoscaling, monitoring | Platform team has SRE capacity or managed Kubernetes |

Use the table as a scorecard. If three or more rows point to self-hosting, start a proof-of-concept with your highest-volume use case.

## What to do next

1. **Audit your current gateway** — Does it abstract the model endpoint, or are product teams calling vendor SDKs directly? If the latter, introduce a thin gateway layer first; it is the prerequisite for any migration.
2. **Build the evaluation harness before you migrate** — Golden datasets, automated metrics and a promotion gate let you prove parity (or superiority) of an open-weight model on *your* tasks, not on public benchmarks.
3. **Run a cost-model comparison** — Model your top five use cases at current and projected volume against reserved GPU pricing (cloud or on-prem). Include engineering ops time; it is not zero.
4. **Map data-residency requirements** — Legal and security teams can enumerate which use cases *must* stay in-VPC. Those become your first self-hosted candidates.
5. **Pilot one high-volume, well-evaluated use case** — Migrate it end-to-end through the gateway, measure the accuracy–latency–cost triangle, and use the results to justify broader rollout.

A & A Labs helps teams design and operate this exact stack — gateway, evaluation harness, self-hosted inference and the operational guardrails that keep answers accurate, source-attributed and affordable.

## Frequently asked questions

### Why did DoorDash move from vendor APIs to self-hosted open-weight models?

DoorDash moved to self-hosted open-weight models to gain control over model versions and quantisation, achieve predictable costs through reserved GPU capacity, keep prompts and completions inside their VPC for data residency and compliance, and reuse a single evaluation harness across model changes.

### What role did the LLM Gateway play in DoorDash's migration?

The LLM Gateway provided a stable /v1/chat/completions-compatible interface that handled request routing, authentication, rate limiting, observability, fallback chains, and prompt-template versioning. It absorbed model-specific differences so product teams could migrate use cases one at a time without rewriting application code.

### How does DoorDash evaluate open-weight models before promoting them?

DoorDash uses a reusable evaluation harness with golden datasets per use case, automated metric suites covering accuracy, hallucination rate, latency percentiles, and cost per 1k tokens, and a promotion gate that blocks model upgrades unless they meet defined thresholds.

### What is DoorDash's decision framework for switching to self-hosted models?

DoorDash uses a five-criterion scorecard comparing vendor APIs and self-hosted open-weight models on model ownership, cost predictability, data residency, evaluation portability, and operational maturity. If three or more criteria favour self-hosting, they recommend starting a proof-of-concept with the highest-volume use case.

### What steps does DoorDash recommend for teams considering a similar migration?

DoorDash recommends auditing the current gateway to ensure it abstracts model endpoints, building an evaluation harness with golden datasets and promotion gates before migrating, running a cost-model comparison for top use cases against reserved GPU pricing, mapping data-residency requirements with legal and security teams, and piloting one high-volume, well-evaluated use case end-to-end.

## Sources

- Presentation: Building GenAI Platform at DoorDash — https://www.infoq.com/presentations/doordash-genai-platform-architecture/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=global

---

Published by A & A Labs · https://aa-labs.co/blog/how-doordash-moved-from-vendor-first-to-open-weight-models-in-its-genai-platform
