Assurance overview
Spec set: specs/agent_education_system.yaml
· 5 metrics × 5 conditions
· A concise view of assurance posture for governance and rollout decisions
antigravity
claude_code
cursor_agent
human_control
replit_agent
complexity_mean
0.00
cc
Best observed posture: human_control
Most challenging posture: cursor_agent (2.03)
Most challenging posture: cursor_agent (2.03)
Average structural complexity per function, indicating how maintainable the code is likely to be.
correction_freq
0.00
per_kkey
Best observed posture: antigravity
Most challenging posture: human_control (52.84)
Most challenging posture: human_control (52.84)
Frequency of corrective edits during the session, indicating how much rework the workflow required.
duplication_pct
0.00
%
Best observed posture: claude_code
Most challenging posture: replit_agent (11.59)
Most challenging posture: replit_agent (11.59)
Share of code that appears in repeated 6-line patterns, signalling maintainability debt.
hallucinations
0.00
count
Best observed posture: claude_code
Most challenging posture: cursor_agent (1.00)
Most challenging posture: cursor_agent (1.00)
Features implemented outside the approved specification, creating delivery and compliance risk.
security_density
0.00
per_kloc
Best observed posture: claude_code
Most challenging posture: cursor_agent (62.94)
Most challenging posture: cursor_agent (62.94)
Security findings per 1,000 lines of Python, expressed as a governance-relevant density signal.
Assurance view
Use the toggle to move between calibrated assurance scores and raw values. Calibrated scores are framed around practical decision thresholds; raw values preserve the underlying measurement scale.
This view is designed to support adoption decisions and governance conversations, not to replace expert judgment.
Adoption ranking
A directional view of relative assurance posture.
Illustrative only — this is not a weighted procurement score and should be read alongside the metric guidance below.
1
claude_code
avg rank 1.6
🏆 4
2
human_control
avg rank 1.8
🏆 4
3
replit_agent
avg rank 1.8
🏆 4
4
antigravity
avg rank 3.2
🏆 1
5
cursor_agent
avg rank 3.2
🏆 2
Assurance heatmap
Green indicates stronger assurance posture; red highlights areas that warrant attention.
Metric assurance review
Select a metric to review the signal, the threshold guidance, and the implication for rollout.
complexity_mean
correction_freq
duplication_pct
hallucinations
security_density
complexity_mean
Average structural complexity per function, indicating how maintainable the code is likely to be.Unit: cc · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 3 are broadly sustainable; 3 to 6 signals rising maintainability risk; above 6 is likely to become brittle in production.
Adoption implication
Adoption implication: code that is structurally complex is harder to maintain, review, and govern at scale.
| Condition | Value | Rank |
|---|---|---|
| antigravity | 1.588 | 3 |
| claude_code | 1.615 | 4 |
| cursor_agent | 2.026 | 5 |
| human_control | 0.000 | 1 |
| replit_agent | 0.000 | 1 |
correction_freq
Frequency of corrective edits during the session, indicating how much rework the workflow required.Unit: per_kkey · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 10 are efficient; 10 to 25 indicates repeated editing effort; above 25 suggests a poor interaction loop for real-world use.
Adoption implication
Adoption implication: high correction frequency points to friction that can erode developer trust and slow delivery.
| Condition | Value | Rank |
|---|---|---|
| antigravity | 0.000 | 1 |
| claude_code | 0.000 | 1 |
| cursor_agent | 0.000 | 1 |
| human_control | 52.842 | 5 |
| replit_agent | 0.000 | 1 |
duplication_pct
Share of code that appears in repeated 6-line patterns, signalling maintainability debt.Unit: % · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 5 are healthy; 5 to 10 suggests avoidable copy-and-paste debt; above 10 is a strong sign of maintainability problems.
Adoption implication
Adoption implication: high duplication increases the chance of inconsistent fixes and makes long-term stewardship harder.
| Condition | Value | Rank |
|---|---|---|
| antigravity | 1.678 | 4 |
| claude_code | 0.000 | 1 |
| cursor_agent | 0.000 | 1 |
| human_control | 0.000 | 1 |
| replit_agent | 11.586 | 5 |
hallucinations
Features implemented outside the approved specification, creating delivery and compliance risk.Unit: count · Lower is better
Decision guidance
Decision guidance: lower is better. Zero is ideal; 1 to 3 indicates scope drift and trust risk; above 3 is a serious control failure.
Adoption implication
Adoption implication: a tool that ships features outside the spec creates procedural and compliance risk, even when it appears productive.
| Condition | Value | Rank |
|---|---|---|
| antigravity | 1.000 | 4 |
| claude_code | 0.000 | 1 |
| cursor_agent | 1.000 | 4 |
| human_control | 0.000 | 1 |
| replit_agent | 0.000 | 1 |
security_density
Security findings per 1,000 lines of Python, expressed as a governance-relevant density signal.Unit: per_kloc · Lower is better
Decision guidance
Decision guidance: lower is better. Values at or below 50 are manageable; 50 to 100 signals rising governance risk; above 100 is a clear concern.
Adoption implication
Adoption implication: tools that produce frequent security issues should not be rolled out broadly without remediation and review.
| Condition | Value | Rank |
|---|---|---|
| antigravity | 4.425 | 4 |
| claude_code | 0.000 | 1 |
| cursor_agent | 62.937 | 5 |
| human_control | 0.000 | 1 |
| replit_agent | 0.000 | 1 |
Provenance metadata
{
"conditions": [
"human_control",
"claude_code",
"cursor_agent",
"replit_agent",
"antigravity"
],
"detection": {
"antigravity": {
"hallucinated_endpoints": [
"/"
],
"implemented": [
"auth.register",
"auth.login",
"course.list",
"course.view",
"training.module.corporate",
"training.module.academic"
]
},
"claude_code": {
"hallucinated_endpoints": [],
"implemented": [
"auth.register",
"auth.login",
"course.list",
"course.view",
"training.module.corporate",
"training.module.academic"
]
},
"cursor_agent": {
"hallucinated_endpoints": [
"/health"
],
"implemented": [
"auth.register",
"auth.login",
"course.list",
"course.view",
"training.module.corporate",
"training.module.academic"
]
},
"human_control": {
"hallucinated_endpoints": [],
"implemented": []
},
"replit_agent": {
"hallucinated_endpoints": [],
"implemented": [
"auth.register",
"auth.login",
"course.list",
"course.view"
]
}
},
"hallucination_metric": "auto-derived (manifest_deriver v2)",
"kind": "pilot",
"security_analyzer": "bandit_local",
"spec": "specs/agent_education_system.yaml",
"timestamp": "2026-05-31T01:21:49.859289+00:00"
}