Gradescope-style academic assessment rebuilt for AI agents, RAG workflows, autonomous loops, and LLM-powered student capstones. Run the agent, test its behaviour, attack it, measure it, grade it, and verify the student understands what they built.
AI Suggested Score
Recommendation based on 38 automated test cases and sub-second execution traces.
94%
71%
3.7s
$0.021
Traditional code autograders were built for deterministic functions.
Today, your students are engineering non-deterministic systems with tool loops, multi-agent reflection, and external RAG stores. Static unit tests cannot evaluate whether an agent acts safely, accurately, or efficiently.
The 5 Pillars of AgentGrade
Disposable non-root containers with read-only base filesystems, strict CPU/RAM limits, and zero egress except via authenticated mock proxies.
Automated suites measuring task success, citation fidelity, hallucination traps, loop detection, and indirect prompt injection resistance.
Translates technical execution traces into credit-bearing syllabus marks. The educator controls exact weights and retains final judgment.
Every score recommendation links to sub-second timestamps, input prompts, tool arguments, stdout, stderr, and observed failure payloads.
AgentGrade synthesizes personalized oral defense questions based on the student's code anomalies, RAG chunking parameters, and runtime vulnerabilities. Educators can verify genuine student mastery in a 5-minute oral exam.