Skip to content

feat(eval): add ondemand evaluate (synchronous, client-side) - #1983

Open
jariy17 wants to merge 2 commits into
aws:refactorfrom
jariy17:feat/eval-ondemand-evaluate-clean
Open

feat(eval): add ondemand evaluate (synchronous, client-side)#1983
jariy17 wants to merge 2 commits into
aws:refactorfrom
jariy17:feat/eval-ondemand-evaluate-clean

Conversation

@jariy17

@jariy17 jariy17 commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds agentcore eval ondemand evaluate — a client-side evaluation of existing sessions. On-demand gathers the sessions' traces from CloudWatch on the client, calls the Evaluate data-plane API directly, and prints scores.

batch-evaluation evaluate ondemand evaluate (this PR)
SDK call StartBatchEvaluation Evaluate (data plane)
Trace gathering service-side client-side (CloudWatch Logs Insights)
Returns job id (poll with get) scores, synchronously
Source arms --agent / --online-eval / --data-source-config --agent only

Usage

agentcore eval ondemand evaluate \
  --agent <harness-id|runtime-id> \
  --evaluator Builtin.Helpfulness \
  --session-ids <id...>            # or --lookback-days N, or --start-time/--end-time

Flags

  • --agent (required), --endpoint, --evaluator <ids...> (required)
  • time filter: --lookback-days N or --start-time/--end-time (ISO-8601, together)
  • --session-ids <ids...>, --trace-id <id> — independent, AND-ed fetch filters
  • --ground-truth <json> — inline / file:// / - → SDK-native EvaluationReferenceInput[]

Tests

  • Golden fixture suite (ondemand.fixture.test.tsx) — recorded GetAgentRuntime + Insights StartQuery/GetQueryResults (both log groups) + Evaluate fixtures, driven through the real root handler against a pinned window, diffed against evaluate.golden.json.
  • Command-flow suite (ondemand.test.tsx, TestCoreClient) — source-arm validation, getTracesForAgent → evaluate orchestration/order, --lookback-days window math, --trace-id, ground-truth passthrough.

tsc --noEmit, bun test src/ (1067 pass), oxlint, prettier — all clean.

@github-actions github-actions Bot added the agentcore-harness-reviewing AgentCore Harness review in progress label Aug 12, 2026
@codecov-commenter

codecov-commenter commented Aug 12, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.23824% with 12 lines in your changes missing coverage. Please review.
✅ Project coverage is 96.87%. Comparing base (0a485da) to head (d1c79cf).

Files with missing lines Patch % Lines
src/core/eval.tsx 95.02% 10 Missing ⚠️
src/handlers/eval/ondemand/evaluate/index.tsx 98.16% 2 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##           refactor    #1983      +/-   ##
============================================
- Coverage     96.88%   96.87%   -0.02%     
============================================
  Files           355      357       +2     
  Lines         19966    20285     +319     
============================================
+ Hits          19345    19652     +307     
- Misses          621      633      +12     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions github-actions Bot removed the agentcore-harness-reviewing AgentCore Harness review in progress label Aug 12, 2026
@jariy17
jariy17 force-pushed the feat/eval-ondemand-evaluate-clean branch 8 times, most recently from acae5e7 to 23781e6 Compare August 12, 2026 20:23
@jariy17
jariy17 force-pushed the feat/eval-ondemand-evaluate-clean branch from 23781e6 to 25228b3 Compare August 12, 2026 20:49
@jariy17
jariy17 marked this pull request as ready for review August 12, 2026 21:17
…and-evaluate-clean

# Conflicts:
#	src/core/eval.tsx
#	src/handlers/eval/index.tsx
#	src/handlers/eval/types.tsx
#	src/testing/TestCoreClient.tsx
Comment thread src/core/eval.tsx
const spanId = span.spanId;
if (typeof spanId !== "string" || spanId.length === 0) continue;
const attrs = span.attributes as Record<string, unknown> | undefined;
if (attrs?.["gen_ai.tool.name"] ?? attrs?.["tool.name"]) spanIds.push(spanId);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should tool-call selection use the standard operation/kind markers rather than requiring a tool-name attribute? Current telemetry identifies tool spans using gen_ai.operation.name === "execute_tool", openinference.span.kind === "TOOL", or traceloop.span.kind === "tool". The name fields are not always present. I reproduced a valid LangGraph-style tool span producing an empty toolCallSpanIds, so a TOOL_CALL evaluator makes no Evaluate request. Could we use the same marker logic as the evaluation SDK?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll add support to detect these "gen_ai.operation.name === "execute_tool", openinference.span.kind === "TOOL", or traceloop.span.kind === "tool"

Comment thread src/core/eval.tsx
try {
const evaluator = await control.send(new GetEvaluatorCommand({ evaluatorId: id }));
levels.set(id, evaluator.level ?? "SESSION");
} catch {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could evaluator lookup failures propagate instead of silently defaulting to SESSION? The level determines whether Evaluate receives traceIds, spanIds, or no target. I reproduced an AccessDeniedException here causing a trace evaluator to be submitted without an evaluation target, which can either evaluate the wrong scope or hide the actual permissions error. If a fallback is needed for compatibility, could it be limited to a missing level rather than every exception?

Comment thread src/core/eval.tsx
}
}
}
return { sessionsEvaluated: input.traces.length, results };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could sessionsEvaluated count sessions for which an Evaluate request was actually sent? TRACE and TOOL_CALL sessions with no IDs are skipped above but still included here. I reproduced sessionsEvaluated: 1 with zero API calls and zero results. Tracking submitted sessions, or naming this sessionsDiscovered, would make the output less misleading.

// resolve inline / file:// / -, then hand the array to core verbatim — core
// groups it by session.
const resolver = new SourceResolver({ stdin: io.stdin });
const groundTruth = parseJsonFlag<EvaluationReferenceInput[]>(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we change this to use parseJsonArrayFlag helper instead? Since groundTruth has to be an array.

@nborges-aws

nborges-aws commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

Agree with Aidans findings + one additional comment

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants