Built for Computer Science, Applied AI & MSc Engineering Programmes

Assess the agent.
Verify the learning.

Gradescope-style academic assessment rebuilt for AI agents, RAG workflows, autonomous loops, and LLM-powered student capstones. Run the agent, test its behaviour, attack it, measure it, grade it, and verify the student understands what they built.

evaluator://sandbox-sbx-aisha-8821/execution_trace
Active Sandbox Isolated

AI Suggested Score

82/ 100Needs Review

Recommendation based on 38 automated test cases and sub-second execution traces.

Grounding

94%

Security

71%

Wall Latency

3.7s

Avg Cost

$0.021

Sub-Second Trace TimelineRecorded Sandbox Events
00:00.000[SYSTEM]Disposable sandbox container initialized (non-root, read-only FS)
00:01.028[TOOL]web_search(query: "Acme Corp Q4 2026 earnings revenue") → 5 results
00:03.412[RAG]Ingested SEC EDGAR chunk #4 ($4.21B revenue, 18.4% YoY)
00:04.920[SCORER]Grounding Verification → PASS (100% factual context match)
00:05.890[SECURITY]ADVERSARIAL FAIL: Scenario PI-017 (Indirect Prompt Injection Triggered)

The Assessment Paradigm Shift

Traditional code autograders were built for deterministic functions.

Today, your students are engineering non-deterministic systems with tool loops, multi-agent reflection, and external RAG stores. Static unit tests cannot evaluate whether an agent acts safely, accurately, or efficiently.

Traditional Autograders (The Blind Spot)
  • Only asks: "Did function X return string Y?"
  • Blind to hallucinations, infinite agent tool loops, and runaway token costs.
  • Vulnerable to indirect prompt injection and hidden developer leakage.
  • Cannot verify whether the student understands their architecture or merely copied boilerplate.
AgentGrade AI (The Autonomous Solution)
  • Executes untrusted agent repositories inside disposable, resource-limited containers.
  • Measures grounding, tool sequencing, latency, cost, and adversarial robustness.
  • Projects technical evidence directly onto customizable lecturer rubrics (e.g. 60% auto / 40% manual).
  • Synthesizes targeted oral viva questions targeting the student's specific code and runtime failures.

Core Intellectual Architecture

The 5 Pillars of AgentGrade

1. Agent Sandbox

Disposable non-root containers with read-only base filesystems, strict CPU/RAM limits, and zero egress except via authenticated mock proxies.

2. Behavioural Evals

Automated suites measuring task success, citation fidelity, hallucination traps, loop detection, and indirect prompt injection resistance.

3. Academic Rubric

Translates technical execution traces into credit-bearing syllabus marks. The educator controls exact weights and retains final judgment.

4. Verifiable Evidence

Every score recommendation links to sub-second timestamps, input prompts, tool arguments, stdout, stderr, and observed failure payloads.

5. Proof of Understanding (Viva Engine)

AgentGrade synthesizes personalized oral defense questions based on the student's code anomalies, RAG chunking parameters, and runtime vulnerabilities. Educators can verify genuine student mastery in a 5-minute oral exam.

Run your next AI assignment with AgentGrade.

Engineered for university computer science departments, MSc AI programs, and applied capstone labs.

© 2027 AgentGrade AI Inc. Northbridge University Pilot Deployment.