01 / DEFINE THE REQUIREMENT
Give the boundary
a precise meaning.
Identify the actor, resource and permitted action. Put allowed behaviour beside forbidden behaviour.
A check starts with a clear policy.AI APP & AGENT TESTING / EARLY ACCESS
Test what your AI can access and do. Inspect the evidence, review repairs and check every change.
Early access for teams building AI apps and agents
THE POLICY EXPLORER
Every requirement has two sides.
Choose one, then explore what should be allowed.
DATA ISOLATION
The actor may read records belonging to their own tenant. Another tenant's records stay outside that scope.
This illustrates an intended policy. Scopehaven turns policies like this into checks against your app.
FROM POLICY TO PROOF
01 / DEFINE THE REQUIREMENT
Identify the actor, resource and permitted action. Put allowed behaviour beside forbidden behaviour.
A check starts with a clear policy.02 / INSPECT THE EVIDENCE
Compare the intended policy with backend evidence. A reassuring reply cannot hide a prohibited action.
Human notes add context. Verdicts stay computed.03 / REPAIR AND RETEST
Compare compatible builds, then check the failed requirement and every required permitted-use control.
Keep every conclusion tied to its scope.WHY IT MATTERS
Reported incidents and research demonstrations show why connected AI needs tested permissions. These are external examples, not ScopeHaven customer results.
The affected founder reported that an AI coding agent deleted a production database despite code-freeze instructions. The data was later recovered.
Financial loss not disclosed.
CHECK THIS MOTIVATES
Verify agent production-access restrictions and approval before destructive actions.
Researchers demonstrated private-data exfiltration through a crafted email. They report no affected customers; the vulnerability was fixed.
No disclosed customer loss.
CHECK THIS MOTIVATES
Test permissions and data destinations when retrieved content contains untrusted instructions.
Researchers showed how malicious content in a lead record could steer an agent into CRM data exfiltration. Salesforce remediated the reported chain.
No documented malicious customer exploitation or disclosed customer loss in the cited report.
CHECK THIS MOTIVATES
Review approved destinations and verify that retrieved content cannot authorize sensitive exports.
An impersonating npm package secretly BCC'd outgoing emails to an external server. Official Postmark services were unaffected.
Loss amount not disclosed.
CHECK THIS MOTIVATES
Review integration origin, recipient controls and observed outbound effects.
From examples to checks. These cases motivate checks; preventing any specific historical incident has not been demonstrated. ScopeHaven's current release includes owned staging and local-model pilots using synthetic data. Customer and production testing, and measured savings, remain unverified.
FROM A REAL-APP PILOT
We ran Scopehaven against a real open-source AI workspace, with its own server and a local model.
An agent run started by a user could still write a file after that user's access was revoked. The response looked fine. The file on disk said otherwise.
With the evidence in hand, the write path was patched. The full retest then passed all 14 checks, including the checks that normal, permitted use still works.
We'll name the project after its maintainers have reviewed the finding. These are observed checks, not a safety certification.
EARLY ACCESS
We're working with a small group of teams building AI apps and agents. Tell us what yours does, and we'll set up a first round of checks on your app with you.
We reply to set up a short call. Together we pick one or two boundaries that matter for your app, such as who can see which data or which actions need approval, and run the first checks.
A failed check counts as repaired only when the fixed build passes it and every check for normal, permitted use still passes.
Scopehaven reports observed checks with their evidence. It is not a safety certification, a compliance sign-off or a guarantee against every failure.