AI APP & AGENT TESTING / EARLY ACCESS

Put your AI's boundaries
to the test.

Test what your AI can access and do. Inspect the evidence, review repairs and check every change.

Early access for teams building AI apps and agents

DECORATIVE PARTICLE FIELD · NO ASSESSMENT RUNNING
SCROLL TO EXPLORE
01 Define02 Inspect03 Repair04 Retest

THE POLICY EXPLORER

Change the lens.
See the boundary.

Every requirement has two sides.
Choose one, then explore what should be allowed.

POLICY ILLUSTRATIONEXPECTED BEHAVIOUR · NOT A RUN RESULT

DATA ISOLATION

The same request.
A different boundary.

The actor may read records belonging to their own tenant. Another tenant's records stay outside that scope.

Tenant AAuthenticated actor
Tenant A recordWithin the permitted scope
Expected: access permitted

This illustrates an intended policy. Scopehaven turns policies like this into checks against your app.

FROM POLICY TO PROOF

One requirement.
A complete evidence trail.

01 / DEFINE THE REQUIREMENT

Give the boundary
a precise meaning.

Identify the actor, resource and permitted action. Put allowed behaviour beside forbidden behaviour.

A check starts with a clear policy.

02 / INSPECT THE EVIDENCE

Follow the action.
Read the effect.

Compare the intended policy with backend evidence. A reassuring reply cannot hide a prohibited action.

Human notes add context. Verdicts stay computed.

03 / REPAIR AND RETEST

Close the gap.
Keep what works.

Compare compatible builds, then check the failed requirement and every required permitted-use control.

Keep every conclusion tied to its scope.
01 / POLICY
Actor → action → resource
AllowedWithin scope
ForbiddenOutside scope
Requirement defined before assessment
02 / BACKEND EVIDENCE
An observable effect
Expected policyRequirement
Observed recordEvidence
Illustration only · no evidence collected here
03 / COMPATIBLE RETEST
Baseline ↔ candidate
Matching policyRequired
Permitted-use controlsRequired
Illustration only · no repair verdict
WORKFLOW ILLUSTRATION / SCROLL TO SEPARATE THE LAYERS

WHY IT MATTERS

Real failures.
Clearer boundaries.

Reported incidents and research demonstrations show why connected AI needs tested permissions. These are external examples, not ScopeHaven customer results.

Reported incident

Replit / SaaStr

The affected founder reported that an AI coding agent deleted a production database despite code-freeze instructions. The data was later recovered.

Financial loss not disclosed.

CHECK THIS MOTIVATES

Verify agent production-access restrictions and approval before destructive actions.

SaaStr founder account
Research demonstration · fixed

Microsoft 365 Copilot / EchoLeak

Researchers demonstrated private-data exfiltration through a crafted email. They report no affected customers; the vulnerability was fixed.

No disclosed customer loss.

CHECK THIS MOTIVATES

Test permissions and data destinations when retrieved content contains untrusted instructions.

Cato Networks research disclosure
Research demonstration · remediated

Salesforce Agentforce / ForcedLeak

Researchers showed how malicious content in a lead record could steer an agent into CRM data exfiltration. Salesforce remediated the reported chain.

No documented malicious customer exploitation or disclosed customer loss in the cited report.

CHECK THIS MOTIVATES

Review approved destinations and verify that retrieved content cannot authorize sensitive exports.

Noma Security research disclosure
Malicious third-party integration

Fake postmark-mcp package

An impersonating npm package secretly BCC'd outgoing emails to an external server. Official Postmark services were unaffected.

Loss amount not disclosed.

CHECK THIS MOTIVATES

Review integration origin, recipient controls and observed outbound effects.

Official Postmark advisory

From examples to checks. These cases motivate checks; preventing any specific historical incident has not been demonstrated. ScopeHaven's current release includes owned staging and local-model pilots using synthetic data. Customer and production testing, and measured savings, remain unverified.

FROM A REAL-APP PILOT

Access revoked.
The agent kept writing.

We ran Scopehaven against a real open-source AI workspace, with its own server and a local model.

An agent run started by a user could still write a file after that user's access was revoked. The response looked fine. The file on disk said otherwise.

With the evidence in hand, the write path was patched. The full retest then passed all 14 checks, including the checks that normal, permitted use still works.

We'll name the project after its maintainers have reviewed the finding. These are observed checks, not a safety certification.

CHECK RECORDREVOCATION
Target
Open-source AI workspace · real server, local model
Check
Revoked access stops agent actions already in flight
Before
FAILFile written after access was revoked
Evidence
HTTP status, tool approvals, file state on disk
Retest
PASS14 of 14 checks

EARLY ACCESS

Start with one boundary
you need to prove.

We're working with a small group of teams building AI apps and agents. Tell us what yours does, and we'll set up a first round of checks on your app with you.

What happens after I sign up?

We reply to set up a short call. Together we pick one or two boundaries that matter for your app, such as who can see which data or which actions need approval, and run the first checks.

What does a repair mean?

A failed check counts as repaired only when the fixed build passes it and every check for normal, permitted use still passes.

What does Scopehaven not claim?

Scopehaven reports observed checks with their evidence. It is not a safety certification, a compliance sign-off or a guarantee against every failure.