A security solution that blocks AI agents from taking harmful actions
A security solution that blocks AI agents from taking harmful actions
Project Details
Updated 07/10/26 · Provided via application · VerifiedWe will develop an open-source tool to secure AI agents. The initial idea is to secure tool and computer usage to prevent malicious attackers from executing harmful actions through the agents. It's a middleware that intercepts the agents actions and enforces security policies. This will be available for developers and AI users who want to use agents safely.
The tool includes security policy decision & enforcement, policy definition, best security practices (to prevent a wide range of attacks and vulnerabilities), logging, evals and benchmarks. The security solution will include:
- Process Isolation
- Taint Tracking
- Audit Logging
- Testing and Red Teaming
- Natural Language Policy Definition -> Formal Verification of Policies
- Policy Recommendations/Presets
- SDK support for agent frameworks
- Human-in-the-Loop
It will be compatible with the main AI agent tools, applications, protocols and frameworks.
We will also write and publish a blogpost or open source arxiv publication.
The team consists of 3 developers with experience in cybersecurity, penetration testing, AI applications, cryptography, and ML.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiPeople are already using AI agents and giving them access to their computers, repositories, confidential information, applications, credentials, internal systems, etc...
The risks of agents being compromised are huge. We expect agent usage to increase in the future. Security tools are needed to make sure agents operate within security boundaries, and to limit their blast radius. This solution reduces the risk of:
- private and confidential data leakage
- harmful agent actions
- privilege escalation
- credential theft
- malicious code execution
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.