ChengRang

AQuA Ambient Quality Agent

AI Platforms Open Source

Google open-sourced this ambient quality agent on October 8, 2026. It runs unattended inside your own Google Cloud project beside your agent, sweeping production trajectories from Cloud Trace, Cloud Logging or BigQuery on a schedule, after each deployment or on demand. A five stage pipeline samples up to 1,000 recent sessions, grades each against a nine point checklist of common breakdowns, clusters findings that share a failure mechanism, has a separate model verify each cluster against up to three full transcripts, then tracks surviving clusters in BigQuery as new, recurring or auto-resolved after 14 days unseen. Raw transcripts, source snapshots and tables stay inside your project boundary, and AQuA never sits in the request path or writes back. Root cause analysis cites line ranges against the deploy-time source snapshot and proposes an edit, but never applies one or opens a pull request on its own

Agent ObservabilityProduction DiagnosisTrajectory ReviewGoogle CloudOpen Source
Visit AQuA Ambient Quality Agent

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

AQuA is the ambient quality agent Google open-sourced on October 8, 2026, short for Ambient Quality Agent. It addresses the problem that starts after an agent ships: offline evals pass, online pass-rate monitoring looks fine, and yet nobody can say exactly where real users get stuck.

It works by running unattended beside your agent inside your own Google Cloud project, sweeping a batch of production trajectories from Cloud Trace, Cloud Logging or BigQuery on a schedule. Three boundaries are worth fixing in mind first: raw session transcripts, source snapshots and BigQuery tables stay inside your project boundary, it never sits in the request path, and it never writes back to your agent. It is an observer, not a participant.

The official analogy fits well: think of a junior quality engineer doing the first pass, reading conversations, filtering out noise, and preparing the case file with evidence sessions for whoever is on rotation. It turns raw production traffic into diagnosed, code-anchored insights, then feeds those archived failure transcripts back into your offline inner loop.

Key Features

Use Cases

Pros

Pricing

Released as an open-source project whose components you deploy into your own Google Cloud project. Deployment creates cloud resources such as a BigQuery dataset and a Cloud Run dashboard, and the underlying Google Cloud infrastructure and model calls bill under their own rates. The sampling cap and single-pass default review exist precisely to keep that spend predictable.

Summary

AQuA covers the stretch of agent engineering most often skipped: the time after launch when nobody is watching. Offline evals and online pass-rate monitoring each have a job, but neither answers which step failed in real sessions. A five-stage pipeline turns production trajectories into evidence-backed insight clusters, an independent model verifies them, and the result is anchored to the source snapshot captured at deploy time.

Several design choices are notably restrained. It explicitly stays out of the request path, never writes back, never applies an edit or opens a pull request, and keeps data inside your own boundary. Review intensity comes in two tiers with an economical single-pass default that keeps scheduled sweeps affordable. Those choices make it an assistant rather than a takeover.

It ships as composable open-source building blocks, and Google has said it wants to explore continuous agent quality together with the community, so interfaces and flows are still evolving. It suits teams that already run production agents but lack a way to attribute failures in the wild.

Version History

Category
AI Platforms
Pricing
Open Source
Tags
Agent Observability · Production Diagnosis · Trajectory Review

Related Tools