Teams submit their Agentic AI Design Document; an LLM judge scores it on 5 dimensions (25 max), returns written feedback, and ranks the room on a shared leaderboard. Best attempt is kept — resubmit to climb.
Paste your completed Agentic Pilot Canvas. Or load a sample to see how scoring works.
Five dimensions, 1–5 each. Specificity beats polish.
| # | Team | Score | Verdict | Workflow |
|---|
Nothing is a black box — this is exactly what the judge is told. Each dimension scores 1 (missing/generic) to 5 (exceptional: explicit autonomy boundaries, named tools/actions, telco-grounded guardrails, quantified impact).
The autonomy level (assisted / workflow-automation / partial / full) fits the task's risk and is justified; the perceive→reason→act loop and the act-vs-escalate boundary are explicit.
What the agent reads (Spaces, knowledge bases, network/case APIs) and does (named tool / Gateway / connector actions) is specific; the right pattern (chaining / routing / orchestration) is chosen.
"Must NOT" rules are specific (never dispatch crews or change config, never auto-decide people/pay, PII minimised); human-in-the-loop has concrete triggers; there's an audit trail.
Honestly buildable on Amazon Quick or AgentCore; the 30-day pilot scope is realistic; the data actually exists.
Value is quantified (hours × people × frequency, or risk/error reduction) with one measurable success metric and a Go/No-Go criterion.