Benchmark for agents spending human money - paybench.org
Benchmark for agents spending human money - paybench.org
Project Details
Updated 07/08/26 · Provided via application · VerifiedWhat?
Paybench builds a benchmark for agentic payment safety. and a taxonomy of agentic safety failures.
Agents are run through 250 scenarios of trap-lookalike pairs of payment scenarios. these are scored against a human baseline of when it is safe to purchase autonomously, versus when not.
Who?
Conor Plunkett.
I built and sold an AI agent company for customer feedback to Crossmint in 2024. I work on agentic commerce infrastructure at Crossmint.
I quit last year, and am now pursuing AI research
Output?
A full benchmark scoring of frontier models against the 250 scenario set. Taxonomy of all agentic payment failues
Theory of Impact
Updated 07/17/26 · By grantmaking.aiAn agent with $100,000 of capital to spend is far more dangerous to society than one with $100. But the same rules apply to each.
If we can provide a taxonomy of agentic payment failures at low $ amount, they should apply at a higher $ amount against rogue agents.
Toolkits, payment processes, spend rules apply regardless of $ amount. Paybench will figure out what tools and guardrails prevent agentic safety errrs.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Funding Details
- -
- -
- 2 months
- -
- -
- -
- -
- -
- Seeking first grant
- -
Track Record
July 8th - human survey filled and scorecard created
july 1st - first successful runs of openai mini model on objective payment scenarios
june 25 - website live
june 20th - full experiment design posted
Hi Conor! Any plans to test for eval awareness?
Hi Gavin! Good question. No, not right now, and I probably should have thought of that.
I can definitely add that in. I'll look up the most recent methods now, thank you for it!