What Basecamp is building
Most biological datasets are not merely small; they are uneven. According to Basecamp’s research on its BaseData collection, about 68% of the sequences in the Sequence Read Archive come from just five species, with humans accounting for roughly 54%.
That concentration makes sense for medical research, but it gives a model a narrow view of biology. Basecamp compares it to training a language model only on newspaper articles from 1975: useful information, but a poor representation of the world.
The company is collecting samples with local research partners from:
The network covers more than 30 countries and all seven continents. Microsoft counted 208 participating organizations in 31 countries.
Basecamp Research works with researchers in dozens of countries to collect samples on site. | Image: Basecamp Research
Source: the-decoder.com
The dataset contained about 15 trillion tokens at the time of the interview. That is comparable to the text corpora used to train models such as Claude and GPT, but the token here is a building block of DNA rather than a word.
The first generation of EDEN was trained on 9.7 trillion DNA building blocks from more than one million recently sequenced species. Basecamp trained it with Microsoft researchers on Azure, using computing resources reportedly comparable to GPT-4. OpenAI has never officially disclosed how much compute GPT-4 required.
The dataset is supposed to grow roughly 100-fold over the next year and a half, surpassing one quadrillion tokens. The planned Trillion Gene Atlas, announced in March, is intended to provide genetic data at the scale of one trillion genes. Basecamp, Anthropic, Nvidia, PacBio and Ultima Genomics are working on it.
Basecamp now describes the atlas as the foundation for future EDEN training. It has not said how far that expansion has progressed.
The company also records the chemical, physical and ecological conditions around each sample. Long, continuous DNA sequences are preferred because they preserve information about which genes sit next to each other and may work together. Short, isolated fragments often lose that context.
Source: the-decoder.com
From microbial competition to drug candidates
Basecamp’s biological thesis is that evolution connects environmental samples to medicine. A bacterium adapting to a hot spring and a cancer cell spreading through the body face different conditions, but both are shaped by selection.

This figure from a 2024 paper compares the protein diversity captured in public databases with Basecamp's holdings at the time. | Image: Vince, Gowers, and McGibbon, GEN Biotechnology, CC BY 4.0
Source: the-decoder.com
Microbial competition offers a more direct route to therapies. Bacteria sometimes produce molecules that suppress or destroy rival species, giving them access to nutrients and territory. Those interactions happen across countless ecosystems, particularly where resources are scarce. Some of the resulting molecules can function as antibiotics.
Bacteriophages provide another example. These viruses infect bacteria and inject their own DNA into bacterial cells. Interactions between phages and bacteria helped produce tools that can be adapted for gene editing.
Basecamp wants EDEN to learn from this broad set of evolutionary solutions rather than simply rediscover known compounds. The model is meant to design new candidates from the patterns it absorbs.
The early evidence is promising but narrow. In an as-yet-unpeer-reviewed paper on the EDEN model family, 97% of a small, selected set of antimicrobial peptides showed activity in laboratory tests. That figure applies only to the group chosen for testing, not to every compound the model can generate.
The company also reported an animal experiment involving EDEN-7. In mice infected with bacteria resistant to multiple drugs, the candidate reportedly performed about as well as an antibiotic of last resort.
Basecamp says EDEN designed the candidate directly, without preliminary rounds of refinement. The work was conducted with University of Pennsylvania scientists led by César de la Fuente.
The model has been used for other biological designs as well. Microsoft reported that EDEN created a synthetic gut microbiome containing about 9,000 bacterial species. Jill Banfield, a microbiome researcher at the University of California, Berkeley, helped set strict evaluation criteria and co-authored related research.
That independent skepticism matters. A model that generates plausible sequences is not the same thing as a system that produces useful medicine.
Source: the-decoder.com
The harder problem starts after design
Basecamp is also using EDEN to design biological tools for inserting large sections of DNA into selected locations in the human genome. Its focus is on large serine recombinases, enzymes that can join DNA segments.
The intended use is to insert a healthy copy of a gene without correcting every mutation separately. That could be useful for inherited diseases caused by different mutations in different patients.
One possible application is CAR-T therapy. These are genetically modified immune cells designed to recognize and destroy cancer cells. In a January announcement, Basecamp said modified human primary T cells destroyed more than 90% of tumor cells in laboratory tests.
The new funding is meant to advance therapies that make such changes inside the body, rather than modifying cells through a complex manufacturing process. Basecamp says that process can cost hundreds of thousands of dollars per patient.
The difficulty is that successful genetic modification in a dish does not explain how to deliver the necessary components to the right cells in a human body. Patrick Hsu, co-founder of Arc Institute, identified delivery, possible toxic effects and cellular immune responses as unresolved problems in a May 2025 Financial Times article.
Basecamp’s published pipeline shows how much remains before clinical trials:
As of September 2026, none of the company’s six programs had moved beyond lead-candidate optimization. Four programs were at that stage; solid-tumor CAR-T therapy and peptides for diabetes remained in early research.
The company has not said which program will reach human testing first or when that might happen. It also has not disclosed preclinical results, independent verification or comparisons with other AI systems.
Basecamp's published therapy pipeline shows none of its six programs has left lead optimization yet. | Image: Basecamp Research
Source: the-decoder.com
My read is that EDEN has demonstrated biological activity, not yet a therapeutic platform. The distinction is substantial. A molecule that works against bacteria in mice is an important result, but it does not establish safety, delivery or efficacy in people. The more revealing milestone will be whether Basecamp can move one of these programs through the long sequence of preclinical and clinical work that follows model generation.
Source: the-decoder.com
Scaling data is only half the strategy
Basecamp is not treating environmental data as a replacement for experiments. Its Boston laboratory runs, among other projects, experiments with human T cells. These studies may produce only hundreds or thousands of data points, but they answer focused questions about whether a particular genetic change works in a particular cell type.
The company uses those results for further training and reinforcement learning. According to technical director Philip Lorenz, one-third of Basecamp’s GPUs is currently reserved for reinforcement-learning experiments.
That allocation signals a shift away from simply collecting more DNA. Basecamp still says its dataset has grown by roughly tenfold each year and that each new training run improves its metrics. But it is also trying to steer a broadly trained model toward specific biological tasks.
Lorenz uses perplexity to measure how well a model predicts the next symbol in a sequence. Lower is better. He says Basecamp’s data could save hundreds of thousands to millions of GPU hours to reach a perplexity of 2, while an extrapolation for perplexity 1.5 suggests a difference of billions of GPU hours. Those are projections, not measured results.
Technical scores do not always predict downstream usefulness. StripedHyena, the architecture behind Arc Institute’s Evo model, sometimes achieved lower perplexity than Llama, but Llama performed better on subsequent biological tasks. Basecamp therefore selected Llama.
That is the more important lesson in these comparisons: a better sequence predictor is not automatically a better drug designer. Basecamp says it checks training checkpoints against biological tasks, including prediction of a possible immune response to a molecule. It has not provided specific numbers for that capability.
The company also gives synthetic data a low priority. Lorenz argues that synthetic sequences mostly add points within already known parts of sequence space rather than revealing genuinely new regions. Internal tests sometimes produced functional proteins from synthetic-data-trained models, but did not show real generalization to new tasks.
The data bargain behind the model
Basecamp does not plan to release its models openly. It cites safety requirements and agreements with data-contributing partners. The company says many partners prefer a commercial model that shares revenue with them to open-source software whose use cannot be monitored.
The company also removes data from some viruses and checks generated sequences against databases of known pathogens. Those safeguards may reduce misuse, but they do not prove that dangerous outputs have been fully excluded.
The commercial model depends on cooperation from the places where samples are collected. Lorenz says Basecamp obtains informed consent from landowners and local and national authorities before collecting samples. He also says every token used to pretrain EDEN can be traced to a precise geographic origin and a specific consent agreement.
Basecamp hires local scientists and invests in sequencing equipment and laboratories. Its BaseData research describes licensing payments distributed according to each partner’s share of the training data. By the end of 2024, the company had paid 52 recipients in 19 countries; the fastest first payment came nine months after sample collection.
The split remains contested. The Financial Times reported in 2025 that payments represented 1% of revenue. Researcher Jim Thomas said the traceability was valuable but argued that the share was too small because the data increases the company’s value while the communities supplying it do not receive a fair proportion. Aurélie Dingom of Cameroon’s environment ministry also said a larger share would be fair.
I think this arrangement is not a side issue. Basecamp’s model depends on access to biodiversity that public genomic databases have systematically underrepresented. If local partners are expected to keep expanding that collection, consent and revenue-sharing need to function as durable infrastructure, not simply as evidence of responsible sourcing.
The $140 million round gives Basecamp room to expand the atlas, train larger models and advance six therapeutic programs. But the company’s central tension is now clear: its strongest advantage may be the breadth and provenance of its data, while its biggest scientific risk lies in proving that evolutionary patterns can survive the journey from a remote ecosystem to a safe treatment in a person.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X