🏠 Home
🎬 Simulated judge — scoring runs in your browser for illustration. The live event uses the real Claude on Amazon Bedrock judge (Lambda + shared leaderboard). Same 5-criteria rubric. ← Back to the plan
Telkomsel Cloud ZoneAI & Security Enablement powered by AWS

Agent Design Judge — Live Scoring & Leaderboard

Teams submit their Agentic AI Design Document; an LLM judge scores it on 5 dimensions (25 max), returns written feedback, and ranks the room on a shared leaderboard. Best attempt is kept — resubmit to climb.

⚖️ 5 criteria · 25 max 🤖 Claude on Bedrock (live) · temp 0.1 🏆 Shared leaderboard 📡 Telkomsel · telco use cases

📝 Submit an Agent Design

Paste your completed Agentic Pilot Canvas. Or load a sample to see how scoring works.

Team name
Agentic AI design document (canvas)

📊 Score & Feedback

Five dimensions, 1–5 each. Specificity beats polish.

⚖️
Submit a design (or load a sample) to see the live score, per-dimension breakdown, verdict, and written feedback.

🏆 Room Leaderboard

#TeamScoreVerdictWorkflow

🔍 How the judge scores

Nothing is a black box — this is exactly what the judge is told. Each dimension scores 1 (missing/generic) to 5 (exceptional: explicit autonomy boundaries, named tools/actions, telco-grounded guardrails, quantified impact).

01 · MAX 5

Autonomy Design

The autonomy level (assisted / workflow-automation / partial / full) fits the task's risk and is justified; the perceive→reason→act loop and the act-vs-escalate boundary are explicit.

02 · MAX 5

Tools & Actions

What the agent reads (Spaces, knowledge bases, network/case APIs) and does (named tool / Gateway / connector actions) is specific; the right pattern (chaining / routing / orchestration) is chosen.

03 · MAX 5

Guardrails & Oversight

"Must NOT" rules are specific (never dispatch crews or change config, never auto-decide people/pay, PII minimised); human-in-the-loop has concrete triggers; there's an audit trail.

04 · MAX 5

Feasibility

Honestly buildable on Amazon Quick or AgentCore; the 30-day pilot scope is realistic; the data actually exists.

05 · MAX 5

Business Impact

Value is quantified (hours × people × frequency, or risk/error reduction) with one measurable success metric and a Go/No-Go criterion.

Total → verdict: 🥇 Pilot-Ready (21–25) · 🥈 Strong Concept (16–20) · 🥉 Good Start (11–15) · 🔄 Needs Refinement (<11). The live judge uses Claude Sonnet 4 on Amazon Bedrock at temperature 0.1 for consistency, keeps your best attempt, and writes to a shared S3 leaderboard. This page reproduces the same rubric offline for illustration.
AWS AWS
Agent Design Judge · a simulated Telkomsel Cloud Zone experience · powered by AWS, delivered with AWS Training & Certification
Illustrative only — scoring is heuristic in this demo; the event uses the real Bedrock judge. · Back to the Idea-to-Execution plan