Most teams have no idea what to evaluate — and that's the real problem. AgentScore watches your agent in shadow mode, identifies what matters based on how it actually behaves, and builds your evaluation design from observed sessions.
Evaluating an AI agent isn't like writing unit tests. There's no obvious spec to derive from, no single failure mode to cover first. Most teams either skip evals entirely or write ones that miss what actually goes wrong.
Connect your agent in shadow mode. AgentScore observes real sessions — which tools it calls, where it hesitates, what patterns emerge. After 20 sessions it surfaces a Measurement Recommendation: specific evaluation dimensions and calibration scenarios matched to your agent's actual behaviour. You review it and confirm. Scoring only begins after a human approves.
No eval expertise required. AgentScore observes first, then tells you what to measure — based on how your agent actually behaves.
AgentScore handles the hard parts — eval design, scoring, regression tracking, and attribution — so practitioners can focus on the domain work they actually know.
Each scoring profile activates the dimensions relevant to the agent type. All 11 are available; none are required. Profiles are versioned so history stays comparable.
AgentScore is built around real decisions — not abstract metrics. Here's how it plays out in practice.
The questions that come up most when teams look at AgentScore for the first time.
Phase 1 is live for internal Tricentis teams. Connect your agent, and AgentScore will tell you what to evaluate — based on what it observes. No eval expertise required to get started.
Request access
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.