Archestra has released OpenAPPA, an open-source security engine built to stop data exfiltration caused by prompt injection or model hallucination. The tool runs completely outside the agent prompt and execution loop. Its configuration maps out data sources, audiences, trust levels, and authorities alongside deterministic security enforcement rules.
OpenAPPA blocks attacks while maintaining task utility
The team behind OpenAPPA reports zero successful attacks during tests on the Bench-Corp and AgentThreatBench security benchmarks. By comparison, Claude Code auto mode yielded a 10 percent attack success rate, while Microsoft FIDES permitted 31 percent. OpenAPPA achieved an 89 percent task completion rate on these tests, while Microsoft FIDES completed 41 percent.
The project addresses the limitations of stochastic policy enforcement approaches. Probabilistic judges often top out at 99.3 percent effectiveness, which leaves room for breaches at scale, and classifiers can themselves be prompt injected. OpenAPPA implements an Agentic Permissions Policy Algebra developed by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy.
Deterministic rules control data flow across tool calls
Policies run through a pluggable engine using a single appa.toml configuration file. The engine jointly labels and monitors audiences and trust levels using lattice algebra. Each tool contract defines required audience and trust levels, resulting restrictions on returned data, and an audit trail of actions.

OpenAPPA includes explicit recovery semantics when agents attempt illegal actions. Sanitizers can strip personally identifiable information, authorities can route requests for human approval, and disposable child branches isolate unvetted data reads in transient subagent environments.
Bench-Corp tests multi-step enterprise workflows, and AgentThreatBench operationalizes the OWASP Top 10 for Agentic Applications. OpenAPPA is currently available as a preview.



