scry.io: 300 TB databases making the internet hyper-traversable to research agents
Giving people and their prosocial agents ergonomic SQL query power over the internet, with structured judgement kernels to organize the information across interpretable high-dimensional axes.
Giving people and their prosocial agents ergonomic SQL query power over the internet, with structured judgement kernels to organize the information across interpretable high-dimensional axes.
Project Details
Updated 07/30/26 · Edited by orgAt scry.io, I am giving people raw readonly Structured Query Language access to hundreds of terabytes of well-prepared internet data. I am ingesting much of the internet, over 1M records per second, and having them end up cleaned, dynamically embedded, and well-indexed in performant analytical databases that I give people safe, fair, ergonomic access to. I own significant hardware for these purposes, including over 400 TB of NVMe storage, so I am not under shallower corporate survival constraints like I'd be if renting a $20,000/month (and worse) server from cloud providers, but instead can design my APIs towards longer-time horizon, rarer, more differentiated purposes.
I have had over $1000 in revenue, and after a relaunch soon since significantly updating my hardware and database technology and completely rewriting my software stack, I expect this project to become profitable in the next couple months, while also becoming a public-benefit project backbone to many EA and AI-safety adjacent initiatives.
An important thing to note is the flexibility I have when working with data of this scale. When someone working in AI safety needs 5B quality text-embedding pairs for their research, it becomes trivial for them to explore exactly what they want and for me to upload it to their S3 bucket. When someone wants to mine for phenomenology reports on reddit historical archives, or basically do a vast assortment of internet text traversal operations, it becomes something people can just do by talking to their agent, and that I can just grant them credits, or they can pay 100-1000x cheaper pricing than Amazon Athena over Common Crawl.
Ultimately I intend to ship a robust enough API and service than many tools and startups can simply rely upon Scry as a research substrate. Epistemic infra operations requires agents, skills, and big data, and non-enterprise, solo people have struggled to have genuinely ergonomic, affordable access to big data affordances until now.
Theory of Impact
Updated 07/30/26 · By grantmaking.aiAI x-risk is fueled by ignorance, sickness, and ecological weakness. Overly-siloed information landscapes make it difficult to explore curiosities, unearth key insights, and unblock people doing critical things. When people have new information traversal tools, it allows them to think differently. Exhaustive search, like combing through every GitHub and Gitlab readme and AGENTS.md, or every word of 100M academic papers, with agents able to lexically permute many search ideas, make it much harder to miss certain wordings and ideas. Negative evidence becomes a first-class primitive, where people can update more when they don't find something. It becomes trivial to notice people who work on advanced technical safety relevant topics, are publicly posting that they are struggling to afford a basic token subscription. It also will be easy to share these queries to better highlight the way the world is actually shaped. And with my integration of structured judgement work https://github.com/XyraSinclair/cardinal-harness into Scry, it'll become much easier to say "when Claude Fable 5 assessed all of X proposals with this robust structured judgement kernel by these interpretable attributes, it systematically thought they were ordered in this way".
People
Updated 07/30/26 · By grantmaking.aiTeam Member
Funding Details
- -
- -
- 10 years
- -
- -
- -
- -
- -
- seeking first grant
- -
I just want to add that ergonomic access to big data genuinely has gravity. People just have to extrapolate and understand that IT IS meaningful to have hundreds of terabytes of the most thoughtful writing and work and metadata about it, in a database that agents have low-friction access to. And that people just have to understand that they will experience the benefits of scry and scry links to query results sooner than later. I'm doing the deep things well, this tool already trivially allows people to research countless hypotheses ergonomically that they never could, like if people who mention certain health interventions have more diverse writing as calculated by embedding distances.