Overview
AQuA is the ambient quality agent Google open-sourced on October 8, 2026, short for Ambient Quality Agent. It addresses the problem that starts after an agent ships: offline evals pass, online pass-rate monitoring looks fine, and yet nobody can say exactly where real users get stuck.
It works by running unattended beside your agent inside your own Google Cloud project, sweeping a batch of production trajectories from Cloud Trace, Cloud Logging or BigQuery on a schedule. Three boundaries are worth fixing in mind first: raw session transcripts, source snapshots and BigQuery tables stay inside your project boundary, it never sits in the request path, and it never writes back to your agent. It is an observer, not a participant.
The official analogy fits well: think of a junior quality engineer doing the first pass, reading conversations, filtering out noise, and preparing the case file with evidence sessions for whoever is on rotation. It turns raw production traffic into diagnosed, code-anchored insights, then feeds those archived failure transcripts back into your offline inner loop.
Key Features
- Five-stage sweep pipeline: Each run samples up to 1,000 recent sessions to keep cost predictable, grades each against a nine-point checklist of common breakdowns, clusters findings that share a failure mechanism, has a separate model check each candidate cluster against up to three full transcripts and discard unsupported ones, then matches survivors in BigQuery as new, recurring or auto-resolved after 14 days unseen
- Data stays inside your project boundary: Deployed in your own Google Cloud project, raw session transcripts, source snapshots and BigQuery tables remain within your project boundary throughout. AQuA never sits in the request path and never writes back to your agent
- Steering review with goal.md: Write a plain-English developer goal into goal.md and it is appended to every review prompt, pulling review toward your domain invariants and suppressing stylistic noise. An eval_config.yaml defines deterministic Python custom metrics that run alongside the judge to trend pass rates
- Root cause anchored to the source snapshot: Triggered from the dashboard chat or agents-cli aqua run, it reads failing trajectories against the immutable source snapshot captured at deploy time. When the defect is in your repo it cites path and line range and proposes an edit anchored to the snapshot lines; when the fault lies upstream, in a handoff or in a retrieved payload it attributes the failure to that step with no code diff. It never applies an edit or opens a pull request on its own
- Two levels of review intensity: By default it uses the single-pass session_review judge, one model call per session, keeping scheduled sweeps economical while producing the actual and expected diffs that drive clustering. You can instead opt in to the Gemini platform managed trajectory AutoRaters, which run dedicated per-metric evaluators for task_success, tool_use_quality and trajectory_quality
- Three commands to deploy: agents-cli extension add, infra single-project and deploy attach AQuA to your project, setting up the BigQuery dataset, a Cloud Run dashboard behind Identity-Aware Proxy, and an immutable snapshot of your source tree in Cloud Storage keyed by deployment revision
Use Cases
- [object Object]
- [object Object]
- [object Object]
- [object Object]
- [object Object]
Pros
- Open sourced as composable building blocks, so you can run it in your own project, adapt it to your stack and help shape where it goes
- Raw production data stays inside your project boundary, stays out of the request path and never writes back, so production is untouched
- The pipeline clusters and independently verifies findings, returning evidence-backed insight clusters rather than single alerts
- Root cause analysis anchors to the immutable source snapshot captured at deploy time and cites specific line ranges, so advice never targets code that has already changed
- It distinguishes in-repo defects from upstream or payload failures, avoiding wild goose chases on external problems
- A 1,000 session sampling cap plus single-pass review by default keeps run cost predictable
- Insights are tracked across runs as new, recurring or auto-resolved after 14 days, which suits long-term trend watching
Pricing
Released as an open-source project whose components you deploy into your own Google Cloud project. Deployment creates cloud resources such as a BigQuery dataset and a Cloud Run dashboard, and the underlying Google Cloud infrastructure and model calls bill under their own rates. The sampling cap and single-pass default review exist precisely to keep that spend predictable.
Summary
AQuA covers the stretch of agent engineering most often skipped: the time after launch when nobody is watching. Offline evals and online pass-rate monitoring each have a job, but neither answers which step failed in real sessions. A five-stage pipeline turns production trajectories into evidence-backed insight clusters, an independent model verifies them, and the result is anchored to the source snapshot captured at deploy time.
Several design choices are notably restrained. It explicitly stays out of the request path, never writes back, never applies an edit or opens a pull request, and keeps data inside your own boundary. Review intensity comes in two tiers with an economical single-pass default that keeps scheduled sweeps affordable. Those choices make it an assistant rather than a takeover.
It ships as composable open-source building blocks, and Google has said it wants to explore continuous agent quality together with the community, so interfaces and flows are still evolving. It suits teams that already run production agents but lack a way to attribute failures in the wild.
Version History
- AQuA released as open-source building blocks (2026-10-08): Google published and open-sourced AQuA, the Ambient Quality Agent, which runs unattended beside your agent in your own Google Cloud project and sweeps production trajectories from Cloud Trace, Cloud Logging or BigQuery on a schedule, after each deployment or on demand. The pipeline samples up to 1,000 sessions, grades each against a nine-point checklist, clusters findings sharing a failure mechanism, has a separate model verify each cluster against up to three full transcripts, then tracks survivors in BigQuery as new, recurring or auto-resolved after 14 days unseen. Root cause analysis reads failing trajectories against the immutable deploy-time source snapshot, citing line ranges to propose edits or attributing the failure externally, and never applies changes or opens pull requests on its own. Review is steered through goal.md and eval_config.yaml, defaults to the single-pass session_review judge, and can opt into Gemini platform managed AutoRaters