The Information Company
Why
People die from diseases we should be able to understand and eventually cure. Extending healthy human life is one of the highest-leverage goals humanity can pursue, yet biology remains constrained by the speed and throughput of scientific discovery.
We believe curing most diseases will require intelligence beyond what human researchers alone can provide. Transformers gave humanity a path to scale scientific reasoning far faster than we can scale the number of trained scientists. But intelligence is only useful when it can learn from reality.
AI has consumed much of the high-quality information available on the public internet. Its largest remaining blind spots are domains where data is private, non-digital, expensive, poorly structured, or does not yet exist. Experimental biology has all five characteristics.
The next bottleneck in AI will not only be compute. It will be access to proprietary, high-quality data generated from the physical world.
Data centers for scientific intelligence
Frontier models are rapidly improving at reasoning, but they cannot become capable experimental scientists by learning from papers alone. Experimental work is largely private, non-digital, poorly structured, and full of tacit knowledge that disappears in failed runs, undocumented adjustments, and unpublished results.
The Information Company is building data centers for scientific intelligence in the form of modular autonomous laboratories in shipping containers, where AI agents can design, run, observe, and debug real experiments. We capture the resulting experimental agent trajectories: the goal and starting context, protocol drafts, tool calls, instrument states, raw observations, errors, interventions, revisions, and final outcomes. These trajectories – rather than only assay endpoints – form proprietary training and evaluation data for frontier AI labs and companies building AI scientists.
A failed experiment with complete context can be more valuable for training than a successful experiment reported as a single endpoint. By capturing every decision, action, observation, error, correction, and result, we produce datasets that do not exist anywhere else.
Today, scientific progress scales with scarce trained personnel: every additional research program requires more scientists and PhDs. Based on current AI progress, we expect that within approximately three years, human scientific reasoning will stop being the primary constraint on how much research can be attempted. The bottleneck will shift to experimental capacity – the number of automated instruments and facilities available for models to test hypotheses in the physical world.
AI labs already spend hundreds of billions of dollars on compute infrastructure. As public datasets are exhausted and model capabilities converge, a growing share of that investment will move toward proprietary data production.
The AI data market grew from less than $1 billion three years ago to approximately $10 billion today and is on track to exceed $100 billion by 2030. The rapid growth of companies such as Mercor, Handshake, Micro1, and MicroAGI demonstrates that frontier labs will pay for difficult-to-generate training data that improves model performance.
Biology is our entry point. After year one, we plan to expand into chemistry and materials science so frontier models can learn to operate across many fields of experimental science.
Our long-term vision is a network of data-center-scale research facilities filled with scientific instruments instead of servers. Standardized lab modules let us add capacity far faster than conventional laboratories: rather than constructing each lab as a bespoke project, we can build a factory that ultimately produces and deploys multiple lab modules per day.
We build the facilities end to end, integrating the physical site, instruments, robotics, software, data capture, evaluation, and operations into one system. The first lab modules establish the interfaces and operating model; once validated, the same architecture can be manufactured, deployed, monitored, and improved as one fleet.
Why we are different
Most automated labs are built to discover a drug, improve a proprietary AI scientist, or deliver an assay result. We are building the infrastructure and data layer that teaches frontier models how to perform experimental science itself.
| Company | What it is building | Why we are different |
|---|---|---|
| Lila Sciences | A proprietary AI scientist and autonomous science platform | Its training data is its moat. Selling its best trajectories to OpenAI, Anthropic, or Google DeepMind would mean training the models it needs to outperform. |
| Periodic Labs | AI scientists operating automated laboratories | It is racing to build the scientist, not supplying the picks and shovels to every frontier lab. It is structurally disincentivized to sell its crown-jewel training data. |
| Medra | An AI scientist for autonomous biological research | Like Lila and Periodic, Medra is building a vertically integrated scientist. We are the neutral infrastructure partner that helps many model developers make their own models scientific. |
| Ginkgo Bioworks | Biological foundry services and large-scale data generation | Ginkgo produces assay-specific outputs – screening data, perturbation-response profiles, and similar endpoints. That data can train a biology predictor; it does not teach a general model how to plan, operate, fail, debug, and recover in a lab. |
| Adaptyv Bio | Automated antibody and protein engineering with a CRO-like model | Adaptyv generates narrow, outcome-oriented antibody data. We generate broad agent trajectories designed for general scientific capability. |
| Substrate Bio | Training data for biology models, including binding screens, cell-perturbation profiles, and genomics datasets | Substrate generates biological endpoint data for specialized prediction and design models. We capture complete experimental trajectories that teach general frontier models how to plan, execute, debug, and learn from scientific work. |
None of these companies is purpose-built to capture and sell the complete session of an AI agent doing science. We record the problem, context, decisions, actions, instrument telemetry, failures, corrections, and outcome – not merely the final biological measurement.
This makes us the neutral data and physical-feedback infrastructure layer for scientific AI. Frontier labs need an early, trusted partner that can help them define the interfaces, benchmarks, datasets, and training loops required to turn a general model into a capable scientist. We can become that partner before each lab spends years building fragmented physical infrastructure internally.
This creates a simple competitive game: frontier labs have to buy every useful dataset available because allowing a competitor to train on more high-quality proprietary data risks falling behind. Building internal laboratories does not change that. Owned labs add proprietary data, but they do not replace external supply; to stay competitive, labs must combine both.
Business model, pricing, and revenue potential
We will license curated datasets of experimental trajectories from a proprietary catalog and generate targeted datasets around specific model weaknesses. Frontier AI labs are the initial buyers; pharmaceutical companies provide an additional future customer segment if they become active buyers of proprietary experimental training data. This market is so new that pricing is still value-driven and no stable market price exists.
Initial research partnerships will keep model developers close to the production loop: identify a model weakness, design experiments that expose it, generate the relevant trajectories, and feed the resulting data back into training. This lets us produce datasets against demonstrated capability gaps rather than speculate about what customers need.
Our pricing assumptions are grounded in direct conversations with operators who have built data-selling businesses and experts in life-science data licensing:
- Human-generated data is commonly priced at approximately production cost + 50% on the first sale.
- The same non-exclusive dataset is often licensed to 3–4 buyers; after production cost is recovered, subsequent licenses are nearly pure gross profit.
- Exclusive rights typically cost 3–5× a non-exclusive license.
- Our target is for direct production cost to equal 10–20% of cumulative dataset revenue, implying an 80–90% gross margin at catalog maturity.
Moat
Our durable advantages compound as the infrastructure fleet and dataset catalog grow:
- Unique data: Agent decisions, instrument interactions, failed experiments, recovery behavior, and high-frequency telemetry down to the raw, unprocessed sensor data of scientific instruments are not available in publications.
- Data quality: Standardized execution, complete provenance, and systematic evaluation make every trajectory useful and auditable.
- Throughput: Replicable autonomous lab modules increase experimental volume and shorten the loop between model training and physical feedback.
- Customer relationships and model integration: Close research partnerships let us understand each model’s weaknesses, generate data in direct response, and become deeply integrated into customers’ training and evaluation loops.
- Harness engineering: Better tool interfaces, context management, feedback, graders, and failure recovery make agents running experiments in our facilities more successful at reaching a goal than the underlying model would be on its own.
- Compounding infrastructure: Every experiment improves the software, automation, evaluation systems, protocols, and future lab-module designs that produce the next dataset.
For frontier AI labs, compute will become less of a bottleneck and algorithmic advantages will diffuse. Proprietary access to high-quality experimental reality will become one of their only durable moats.
End state
The training-data business is the wedge. Initially, AI labs pay for the data that improves their models; that revenue finances our infrastructure buildout, while customers bear the much larger cost of frontier-model training.
If we execute, we can ultimately control a substantial share of global experimental capacity. A factory producing standardized lab modules at scale will let us generate scientific data faster, cheaper, and at a volume no conventional laboratory network can match. The resulting feedback loop – more lab modules, more experiments, more data, better models, and still more productive lab modules – compounds with every deployment.
From that position, we can build our own scientific foundation models and direct the infrastructure toward proprietary drug discovery. Running up to 1,000× more experiments than traditional competitors would let us explore more hypotheses, learn faster from failure, and discover drugs on timelines others cannot match. The end state is not merely a data vendor or an automated lab. It is the largest experimental infrastructure network in the world, paired with models trained on a uniquely complete record of physical science.
A company that controls a meaningful share of the world’s experimental capacity can capture value across AI infrastructure, proprietary scientific data, foundation models, and drug discovery – each an enormous market on its own. If we become the default physical infrastructure through which machine intelligence performs science, a trillion-dollar outcome follows from owning the bottleneck beneath many industries rather than winning a single product category.
Roadmap
- 3 months: Build the first lab module, hire the founding team of hardware, software, and ML engineers and founding scientists. Run the first fully autonomous experiments in our lab module.
- 6 months: Open Facility001 and begin the first research partnerships with model developers.
- 12 months: Start building the first data-center-scale research facility. Begin the first paid data and model-training partnerships.
- After year one: Expand into chemistry and materials science.
Example benchmark in the first lab module
The first lab module is designed as a general-purpose biological environment that can eventually support hundreds of experiment types, but the initial proof will deliberately validate one workflow. A candidate benchmark would measure whether a frontier model can express and purify a recombinant protein against explicit yield, purity, and reproducibility targets. The agent would design a plate-based experiment, prepare conditions on a liquid handler, move samples with a robotic arm, monitor cultures by microscopy and plate reader, measure expression with qPCR, separate samples by centrifugation, evaluate purity by SEC, diagnose failures, and launch the next iteration. The exact benchmark will be co-designed with a buyer and may change if research partnerships identify a more valuable capability gap.
The product is not just the final protein or measurement. Each run yields a replayable trajectory containing the experimental objective, available context, protocol versions, liquid-handling and robot commands, instrument telemetry, images and raw measurements, failed steps, recovery actions, human interventions, and outcome-linked quality labels. This supports supervised training, lab-in-the-loop post-training, and capability evaluation.
Indicative lab-module hardware budget
| Component | Low / used | High / new |
|---|---|---|
| 20-foot shipping container | $10k | $20k |
| Robot-accessible cell-culture incubator | $40k | $180k |
| Brooks PreciseFlex plate-handling robot | $20k | $150k |
| Live-cell analysis microscope | $25k | $300k |
| Opentrons automated liquid handler | $15k | $60k |
| Size-exclusion chromatography (SEC) system | $40k | $120k |
| Robot-accessible plate centrifuge | $30k | $100k |
| Multimode microplate reader | $20k | $100k |
| OpenShelf automated plate storage | $15k | $25k |
| Real-time PCR (qPCR) system | $10k | $100k |
| Laboratory refrigerator (2–8 °C) | $1k | $8k |
| Laboratory freezer (−20 °C) | $1k | $12k |
| Ultra-low-temperature freezer (−80 °C) | $8k | $20k |
| Listed hardware subtotal | $235k | $1.20M |
The low case assumes used or surplus hardware that may require refurbishment or automation retrofits; the high case assumes new, premium, automation-ready instruments. Used equipment is generally cheaper and can be faster to acquire when suitable inventory is immediately available, but it can also be slower when we must wait for a specific auction or listing. New equipment is more expensive, but may arrive sooner than the right used system. We prioritize deployment speed above purchase price and will buy new or used equipment based on whichever gets a functioning lab module online first.
The subtotal is an equipment estimate, not an all-in deployment budget. It excludes container fit-out, HVAC and utilities, biosafety systems, electrical and networking work, robotic integration and guarding, software, freight and taxes, consumables, installation, validation, maintenance, and operating staff.
Team
Leon is chief of staff at Dash0, where he helped scale the company from pre-revenue to unicorn in 15 months, building out GTM, hiring, finance, and investor relations. Before Dash0, he went through Y Combinator’s S24 batch and worked in strategy consulting at Bain & Company. He is a former professional esports player and ultramarathon runner.
Claudio built a lab-test platform that processed more than 40 million COVID PCR tests across 5,000 sites as part of a two-person development team. The platform served one million weekly users. He then studied molecular biotechnology at TUM and recently dropped out to get back to building. He built automated workcells for AI-bio companies using different types of lab instruments and is an active contributor to PyLabRobot, the most widely used open-source lab-automation library.
Jialin holds a PhD in precision cancer medicine and brings 10 years of wet- and dry-lab experience across chemistry, materials science, bioengineering, and molecular medicine. She trained and conducted research at Cambridge and Imperial College, has won six hackathons, and is a triathlete.
We are already in talks to hire founding engineers, AI researchers, and scientists to bring together an exceptional early team ready to build this with us.