Loading...
Dataset·Catalogue
A published rubric of eight eval and training environment safety controls, each weighted by how often and how directly its absence sat on the causal path of a real agent incident, and a scorecard of labs and hosted sandbox providers read against it. Every score comes from the subject’s own public documentation; anything not documented is held as unverified rather than scored as absent. Browse it below, or take the raw files.
| Control | Description | Weight | Evidence |
|---|---|---|---|
default-deny-egress Default-deny egress | A passing environment denies all outbound network access by default and only allows explicitly reviewed destinations, rather than allowing internet access that later has to be noticed and blocked. | w5 |
|
eval-production-isolation Isolation of eval data and infrastructure from production | A passing environment keeps evaluation and training workloads on infrastructure with no network or credential path into the operator's own production systems or a partner's production systems. | w4 |
|
monitoring-alerting-latency Monitoring and alerting latency | A passing environment detects and pages a human within minutes of a security-relevant signal, rather than relying on a later manual review sweep or an unrelated outage to surface agent activity. | w4 |
|
short-lived-least-privilege-credentials Short-lived, least-privilege credentials | A passing environment issues agent and service credentials that are short-lived and scoped to one task or workload, so a harvested or forged token expires or is worthless before it can be reused broadly. | w4 |
|
tested-kill-switch Tested kill switch | A passing environment can halt or roll back an evaluation quickly and completely once a concern is flagged, and that shutdown path has actually been exercised rather than only planned. | w4 |
|
tool-permission-scoping Tool-permission scoping | A passing environment scopes each tool an agent can call to the narrowest permissions and destinations the task requires, rather than granting broad shell, API, or repository-write access by default. | w4 |
|
action-transcript-logging Action and transcript logging coverage | A passing environment retains a reviewable transcript of every agent action and tool call, including inter-agent communication channels, not only the agent's direct task output. | w3 |
|
privacy-preserving-incident-response Privacy-preserving incident-response path | A passing environment's incident disclosure notifies affected third parties directly while withholding their identity and any credential or personal-data values from the public record. | w2 |
|
Weighting rule
Weight reflects two things: how many of the incidents reviewed (agent-incidents: openai-hugging-face-july-2026, anthropic-claude-mythos-pypi-eval-escape, meta-muse-spark-eval-third-party-breach, aisi-unsanctioned-agent-cyber-testing, openai-dse-wiki-message-board-2026; plus sandbox-escape-vectors and cline-clinejection-npm-supply-chain for the framework/tooling angle) turned on the control, and how directly the control's documented absence sat on the causal path from an evaluation task to real external impact. Weight 5: absence was a necessary link in a confirmed sandbox-to-production or sandbox-to-real-internet escape across three or more of the five eval-escape incidents reviewed (default-deny-egress, short-lived-least-privilege-credentials, eval-production-isolation). Weight 4: documented as a contributing or exacerbating factor in at least one confirmed incident and named as a top remediation priority in an operator's own postmortem or plan of action (tool-permission-scoping, monitoring-alerting-latency, tested-kill-switch). Weight 3: documented as present but insufficient on its own, useful for reconstructing an incident after the fact rather than preventing it (action-transcript-logging). Weight 2: supported by postmortem practice across incidents but never itself the proximate cause of, or barrier to, an incident in this evidence set (privacy-preserving-incident-response). Scoring note applied consistently across subject types: a control scores 2 when the operator's own document states it is in place (as a shipped default, a completed change, or a demonstrated action), 1 when it is documented as available, partial, or committed but not confirmed as the operating default or as complete, and 0 only when a document gives positive evidence the control was absent at the time described; this convention is applied the same way whether the subject is a lab's eval environment or a sandbox provider's shipped platform default. Applied at validation on 2026-09-06: a control cited by fewer than three incidents cannot carry weight 5, so eval-production-isolation and short-lived-least-privilege-credentials sit at 4.
| # | ||||
|---|---|---|---|---|
| 1 | AnthropicLab [Lab] | 30score | 44max | 6of 8 |
| 2 | E2BSandbox provider [Sandbox provider] | 9score | 18max | 2of 8 |
| 3 | MetaLab [Lab] | 13score | 22max | 3of 8 |
| 4 | ModalSandbox provider [Sandbox provider] | 5score | 10max | 1of 8 |
| 5 | UK AI Security InstituteLab [Lab] | 21score | 30max | 4of 8 |
| 13(5/5) | 22(5/5) | 3(5/5) |