Synthetic Data for AI Agent Evals: Build the Set Before You Have Users
Real production traces are the best eval data you can get, and most teams do not have enough of them yet. Here is the concrete mechanism for generating a synthetic eval set from a handful of real examples, plus the honest failure modes: data that drifts from what real users do,

