feat(eval): add ondemand evaluate (synchronous, client-side) - #1983
feat(eval): add ondemand evaluate (synchronous, client-side)#1983jariy17 wants to merge 2 commits into
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## refactor #1983 +/- ##
============================================
- Coverage 96.88% 96.87% -0.02%
============================================
Files 355 357 +2
Lines 19966 20285 +319
============================================
+ Hits 19345 19652 +307
- Misses 621 633 +12 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
acae5e7 to
23781e6
Compare
23781e6 to
25228b3
Compare
…and-evaluate-clean # Conflicts: # src/core/eval.tsx # src/handlers/eval/index.tsx # src/handlers/eval/types.tsx # src/testing/TestCoreClient.tsx
| const spanId = span.spanId; | ||
| if (typeof spanId !== "string" || spanId.length === 0) continue; | ||
| const attrs = span.attributes as Record<string, unknown> | undefined; | ||
| if (attrs?.["gen_ai.tool.name"] ?? attrs?.["tool.name"]) spanIds.push(spanId); |
There was a problem hiding this comment.
Should tool-call selection use the standard operation/kind markers rather than requiring a tool-name attribute? Current telemetry identifies tool spans using gen_ai.operation.name === "execute_tool", openinference.span.kind === "TOOL", or traceloop.span.kind === "tool". The name fields are not always present. I reproduced a valid LangGraph-style tool span producing an empty toolCallSpanIds, so a TOOL_CALL evaluator makes no Evaluate request. Could we use the same marker logic as the evaluation SDK?
There was a problem hiding this comment.
I'll add support to detect these "gen_ai.operation.name === "execute_tool", openinference.span.kind === "TOOL", or traceloop.span.kind === "tool"
| try { | ||
| const evaluator = await control.send(new GetEvaluatorCommand({ evaluatorId: id })); | ||
| levels.set(id, evaluator.level ?? "SESSION"); | ||
| } catch { |
There was a problem hiding this comment.
Could evaluator lookup failures propagate instead of silently defaulting to SESSION? The level determines whether Evaluate receives traceIds, spanIds, or no target. I reproduced an AccessDeniedException here causing a trace evaluator to be submitted without an evaluation target, which can either evaluate the wrong scope or hide the actual permissions error. If a fallback is needed for compatibility, could it be limited to a missing level rather than every exception?
| } | ||
| } | ||
| } | ||
| return { sessionsEvaluated: input.traces.length, results }; |
There was a problem hiding this comment.
Could sessionsEvaluated count sessions for which an Evaluate request was actually sent? TRACE and TOOL_CALL sessions with no IDs are skipped above but still included here. I reproduced sessionsEvaluated: 1 with zero API calls and zero results. Tracking submitted sessions, or naming this sessionsDiscovered, would make the output less misleading.
| // resolve inline / file:// / -, then hand the array to core verbatim — core | ||
| // groups it by session. | ||
| const resolver = new SourceResolver({ stdin: io.stdin }); | ||
| const groundTruth = parseJsonFlag<EvaluationReferenceInput[]>( |
There was a problem hiding this comment.
Can we change this to use parseJsonArrayFlag helper instead? Since groundTruth has to be an array.
|
Agree with Aidans findings + one additional comment |
Summary
Adds
agentcore eval ondemand evaluate— a client-side evaluation of existing sessions. On-demand gathers the sessions' traces from CloudWatch on the client, calls theEvaluatedata-plane API directly, and prints scores.batch-evaluation evaluateondemand evaluate(this PR)StartBatchEvaluationEvaluate(data plane)get)--agent/--online-eval/--data-source-config--agentonlyUsage
Flags
--agent(required),--endpoint,--evaluator <ids...>(required)--lookback-days Nor--start-time/--end-time(ISO-8601, together)--session-ids <ids...>,--trace-id <id>— independent, AND-ed fetch filters--ground-truth <json>— inline /file:///-→ SDK-nativeEvaluationReferenceInput[]Tests
ondemand.fixture.test.tsx) — recordedGetAgentRuntime+ InsightsStartQuery/GetQueryResults(both log groups) +Evaluatefixtures, driven through the real root handler against a pinned window, diffed againstevaluate.golden.json.ondemand.test.tsx,TestCoreClient) — source-arm validation,getTracesForAgent → evaluateorchestration/order,--lookback-dayswindow math,--trace-id, ground-truth passthrough.tsc --noEmit,bun test src/(1067 pass),oxlint,prettier— all clean.