Investigating how populations become excluded from AI-relevant health datasets
Investigating how populations become excluded from AI-relevant health datasets
Project Details
Updated 07/09/26 · Provided via application · VerifiedProject summary
Many AI systems are trained on datasets that leave out large parts of the world's population. As these systems start influencing healthcare, public health and other important decisions, missing data becomes more than a technical issue. If some populations are not represented, AI systems may simply perform worse for them.
Most work on AI safety looks at models after they have been trained. This project looks further upstream. Before data reaches an AI model, it passes through hospitals and health information systems where it can be lost, remain on paper or never be recorded in a structured way.
During an 8-12 week pilot in healthcare facilities in Kinshasa, Democratic Republic of the Congo, I will map clinical data workflows to understand where routine health data disappears. I will work with participating facilities to implement and evaluate a structured data capture workflow designed to fit existing clinical practice.
The aim is not only to document the problem but to test a practical solution. Outputs will include workflow maps, documentation of where data is lost, an operational data capture configuration for participating facilities and recommendations for improving representation of underserved populations in future health datasets.
What are this project's goals? How will you achieve them?
The goal is to understand why health data from many low- and middle-income countries rarely ends up in the datasets used by AI systems, biosurveillance platforms and epidemiological forecasting models. To achieve this, I will conduct an implementation pilot in healthcare facilities in Kinshasa.
I will map clinical workflows, interview healthcare workers and review existing documentation practices to identify where data becomes incomplete or unusable. Based on these findings, I will introduce a structured data capture workflow and evaluate whether it improves the quality and usability of routine clinical data.
Who is on your team? What's your track record on similar projects?
I am the principal investigator and will lead all aspects of the project, from study design and stakeholder engagement to implementation, analysis and reporting.
My background is in public health and health economics. I have worked in health data management, healthcare information systems, data quality assurance and health economic evaluation. I have also contributed to international development projects, including the United Nations Industrial Development Organization (UNIDO) Annual Report for Madagascar.
Over the last four months I have prepared the implementation framework for the pilot, including the study protocol, workflow mapping tools, interview guides and ethics documentation. These materials are complete and ready for implementation. An overview is available on GitHub :https://github.com/Beeotics/Health-Data-Pilot.
I have also spoken with healthcare professionals and facility leadership in Kinshasa during the planning stage. Several facilities have expressed interest in participating, subject to the necessary approvals.
This project builds on my experience working with health information systems and data quality. Throughout my work, I have repeatedly seen valuable clinical information collected every day but never converted into structured data that can support research, public health or AI development.
What are the most likely causes and outcomes if this project fails?
The most likely risks are operational rather than technical. Potential challenges include limited participation from healthcare facilities, competing demands on healthcare workers' time, difficulties maintaining adoption of structured data capture workflows or delays in obtaining institutional approvals.
If the project fails, the primary consequence would be insufficient evidence to validate the proposed intervention or generate operational observations and lessons learned.
However, even a partially successful implementation would likely generate useful implementation observations into workflow constraints, data quality challenges and barriers to participation in AI-relevant health data systems. The project does not depend on achieving large-scale deployment to produce valuable findings.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiAI systems can only learn from the data available to them. If entire populations are consistently missing from the datasets used to train and evaluate those systems, blind spots become harder to correct as the technology is deployed more widely.
This project looks at one cause of that problem. It examines why routine health data from low- and middle-income countries often never reaches the datasets used to train and evaluate AI systems.
The project is not trying to solve AI alignment. Instead, it focuses on part of the problem that receives much less attention: making sure those datasets better reflect the populations the systems are intended to serve. While this pilot is small in scale, it addresses a bottleneck that affects how health data from many underserved populations enters the broader data ecosystem. Better data alone will not eliminate AI risk, but reducing systematic gaps in representation can help build systems that are more reliable, more robust and less likely to fail for entire populations.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.