The first recipient of the money was a cancer center in Chapel Hill, North Carolina. On September 15, the OpenAI Foundation, the nonprofit that sits atop the company’s structure, announced a new program called Public Data for Health and directed its first $40 million to the Lineberger Comprehensive Cancer Center at the University of North Carolina.
The money has a specific job: collecting the clinical data needed to develop personalized cancer vaccines. It is the opening move in a program whose stated aim is to pay for high-quality scientific datasets, the raw material on which useful medical AI depends and which is often the scarcest thing in the field.
The foundation’s grants are small but pointed. It committed $500,000 to a group called 1Day Sooner to mine the regulatory filings, manufacturing strategies, and safety data of failed biotech companies out of bankruptcy proceedings. A predictive competition run by OpenAdmet is also on the funding list.
The strategy behind the grants is data first. The foundation is not funding a model or a drug. It is funding the datasets that models and drugs will later be trained and tested on, betting that the shortage of good data, not the shortage of compute, is the real bottleneck in applying AI to medicine.
The money itself comes from an unusual place. The OpenAI Foundation holds a 26 percent stake in the company, and that stake is the source of its resources. At a valuation of $1 trillion, the nonprofit could eventually sit on $250 billion in assets, more than the roughly $180 billion held by the Gates Foundation and its affiliated trusts at the end of 2025.
That arithmetic is worth sitting with. The charitable arm of an AI company could become the largest philanthropy in the world, and it would be funded by the very technology whose risks the rest of the field is now debating. The foundation is beginning to define what it will do with that money, and health data is its first answer.
The choice of cancer vaccines is deliberate. Personalized cancer vaccines are an area where clinical data is fragmented, expensive, and held by institutions that rarely share it. Funding the collection of that data at a single research center is a way to build a dataset that no commercial player has an incentive to assemble.
The 1Day Sooner grant points in a different direction. Extracting safety and manufacturing data from bankrupt biotechs is a way to recover knowledge that would otherwise be lost in liquidation, and to make past failures useful to future researchers. It is a small grant with an unusual thesis about where valuable data sits.
Analysts who follow the foundation said the early grants reveal a philosophy. Rather than fund flashy AI projects, the foundation is funding the unglamorous substrate, the datasets and archives, on which everything else depends. That is the kind of spending that compounds but rarely makes headlines.
There are questions the foundation has not answered. How it will govern itself, who will decide where the money goes, and whether its decisions can remain independent of the company’s commercial interests are all unresolved. A nonprofit controlled by a company’s founders is a delicate structure.
The health program is also a hedge on public perception. As OpenAI’s commercial ambitions grow, its nonprofit arm funding cancer research points back to the company’s stated mission, and a way to show that the technology’s proceeds are flowing back to society rather than only to shareholders.
The structure behind the money is unusual by design. OpenAI was founded as a nonprofit, and the capped-profit company that now carries its name sits underneath it. The foundation’s 26 percent stake is the residue of that original structure, and it is the reason a charity could one day control a quarter-trillion dollars in assets.
The dataset bet reflects a hard truth about medical AI. Models are only as good as the data they are trained and tested on, and in medicine that data is locked in institutions, fragmented across formats, and gated by privacy rules. Funding its collection directly is a way to attack the bottleneck that no amount of compute can remove.
Personalized cancer vaccines are a telling first target. They are built from the mutations in a single patient’s tumor, which means every patient generates a unique dataset, and the clinical data needed to validate the approach is expensive and rare. A $40 million grant to one center is the first step in assembling what no company would pay for alone.
Whether the datasets the foundation funds actually accelerate cancer vaccines will take years to know. The program’s first check has been written, and the pattern it sets, paying for data rather than for models, is likely to define the foundation’s next decade.


