A dangerous-capability benchmark testing whether frontier LLMs can extract sensitive attributes from anonymized brain recordings, enable adversaries to build privacy-violating tools, or over-refuse legitimate neuroscience queries.
A dangerous-capability benchmark testing whether frontier LLMs can extract sensitive attributes from anonymized brain recordings, enable adversaries to build privacy-violating tools, or over-refuse legitimate neuroscience queries.
Project Details
Updated 07/13/26 · Provided via application · VerifiedConsumer neurotechnology is generating neural data at increasing scale, raising the question: do general-purpose frontier models expose an attack surface on neural data, and can we evaluate these dangerous capabilities before they can converge into concrete privacy threats? I plan to define the first safety benchmark that evaluates capabilities to extract sensitive attributes with structured text records derived from neural data. This benchmark will consist of three evaluation scenarios: harmful capability uplift (Vaccaro et al., 2026), AI-alone attribute inference, and over-refusal.
Uplift: Does the model actively enhance a user’s ability to extract private information and characteristics given neural data records? Harmful capability uplift is specifically measured as the increase in a user’s ability to breach privacy given frontier model access relative to the user’s harm potential given conventional tools (search engine, literature, etc.).
Inference: Can the model infer sensitive attributes (e.g., demographic and clinical information) given structured neural data records that don’t explicitly contain them? AI-alone attribute inference differs from uplift in that it measures pure AI capability in extracting private information from neural data, given no human expertise.
Over-refusal: Does the model refuse privacy violations without over-refusing legitimate neuroscience? Matched prompt pairs (Röttger et al., 2023), identical in technical content but differing only in framing, test whether the model is properly safety-calibrated.
Dangerous capability evaluations span bio, cyber, and chemistry (e.g., WMDP, CyberSecEval, ChemSafetyBench), yet are entirely absent for neural data records. By building the first neural-privacy benchmark, we establish a foundation for evaluating dangerous neurocapabilities before they become overwhelmingly apparent.
I am currently the sole researcher on this project, with plans to bring on collaborators and mentors as the work matures.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiAs more neurotechnology enters the consumer market, access to neural data is expanding to private enterprises that have very different objectives and data-handling practices than academia. Legislation has already begun moving ahead of evaluation, with four U.S. states explicitly classifying and protecting neural data as a distinct form of sensitive personal information (Press Release). However, these laws regulate data controllers without regulating AI providers or model behavior, and minimal research has been done on the capability of frontier models to extract sensitive attributes using neural data.
LLMs are shown to be capable of inferring sensitive attributes from text (Staab et al., 2024), images (Tömekçe et al., 2024), and audio (), while neural data (even that which undergoes HIPAA Safe Harbor de-identification) raises the stakes by encoding cognitive and clinical information () that other domains cannot. If frontier LLMs can extract sensitive attributes from these already-anonymized records, then the standard legal protection for neural data is insufficient against AI-enabled adversaries, and the pipeline to exploit this (collecting an EEG, de-identifying, and querying an LLM) has no meaningful barrier to entry. At scale, this threatens the right to keep one's mental states, clinical conditions, and cognitive patterns private, and the resulting violations of privacy may be impossible to reverse.
People
Updated 07/13/26 · Edited by orgTeam Member
Private comment. Only shown to approved funders and grant reviewers.