Blog
-
September 15, 2026

Why We Give Our Best Science Away — And Why We Sometimes Don't

Two weeks ago, NVIDIA agreed to pay nearly $13 billion for Hugging Face. Jensen Huang's line in announcing the deal stuck with me: "Open models let startups, businesses, universities and public institutions build on advanced capabilities without training every model from scratch or paying frontier-model prices for every task." A chip company just spent thirteen billion dollars to buy an open-source habit. That should settle, once and for all, an argument I've been having since I left academia almost a decade ago to join a frontier AI lab, then a techbio company, before eventually co-founding Bioptimus: openness is not charity. It's strategy.

I want to explain why, because it's a decision we keep having to re-justify to investors who ask why we'd give away years of R&D, to competitors who don't, and, honestly, to ourselves every time a client chooses our free model over the one we'd rather sell them.

Why we opened H-Optimus

We launched H-Optimus-0, our first foundation model for histology, in 2024 under an Apache 2.0 license: no restrictions, commercial or otherwise. Anyone (an NGO pathology lab in the Amazon, a PhD student in Lyon, a biotech in Boston, a frontier AI lab in SF) could take it, run it, build on it, ship it. When we later released H-Optimus-1 in 2025, trained on more than a million slides from over 800,000 patients, we made a different choice: free for academic and non-commercial research, licensed separately for anyone using it commercially. Two models, two philosophies on the same spectrum, and I think both were right for their moment.

This year we added a third rung. M-Optimus, our newest and most capable model, is the first to fuse pathology images with transcriptomic and spatial data into one multimodal read of a tumor — and it ships under a straight commercial license, full stop. We're not opening it, at least not yet. What we are doing is putting it directly into the hands of a small set of partners working on drug discovery, biomarker identification, and clinical trial design, so we can watch it earn its keep on real questions before we decide how wide to eventually open that door. Call it the same instinct as H-Optimus-1's academic tier, just applied a step earlier in a model's life: controlled exposure before broad release, rather than a locked door indefinitely.

The first reason why we decided to open our models is simple, and it's the only reason that has ever mattered to me: impact. I didn't leave an academic research career to build a better spreadsheet. Patients are our north star, and a model that sits behind a paywall reaches the labs that can afford it, not the labs that need it. Cancer doesn't check a customer's budget before it shows up in a biopsy. If our model helps a pathologist somewhere catch something a human eye would have missed, I don't actually care whether that pathologist works for a top-ten pharma or a public hospital with no procurement budget. Every additional user is a chance for the science to matter to an actual patient — and in this field, that's the only scoreboard I trust.

The second reason is one we talk about less, because it sounds self-interested even though it isn't: we learn more by giving the model away than by keeping it locked up. When thousands of researchers use H-Optimus on data we'll never see, in cancers we didn't train for, on scanners we never tested — we get a map of where the model actually works and where it quietly fails, at a scale no internal validation team could ever replicate. That signal is gold. We're preparing a second post specifically on what two years of open usage has taught us — the surprising places it generalizes, the places it doesn't, and what that says about what "foundation model for biology" should even mean. For now, the short version is: openness isn't just distribution, it's the cheapest, most honest R&D pipeline we have.

The numbers back this up more plainly than any argument I can make. As of this month, the H-Optimus family has been downloaded more than 1.8 million times on Hugging Face — 1.2 million of that H-Optimus-0 alone, with the semi-open H-Optimus-1 already past 400 thousand downloads in barely a year and still climbing. None of that is a vanity metric to me. Every one of those downloads is a lab we didn't have to sell to, using a model we didn't have to support directly, on a question we'll probably never hear about unless it surfaces in a paper (about 150 research papers have already cited our models).

The business case for being open

I understand why this sounds naive in a room full of investors. So let me make the business case, because there is one. An open model is a trust signal in a field — biomedicine — where trust is the actual product. Clinicians and researchers don't adopt a black box from a company they've never heard of; they adopt tools they can inspect, cite, and build their own careers on. Every academic paper written using H-Optimus or M-Optimus is a validation study we didn't have to pay for and a citation trail that makes our commercial offering credible by association. It's also how you recruit: the best people in this field want to publish, and a company that lets them do that is competing for talent on different terms than one that doesn't.

And it's how you build a platform instead of a product. Hugging Face didn't become worth thirteen billion dollars by keeping models proprietary : it became indispensable by being the place where an entire field converges, and NVIDIA just paid a premium to own that convergence point rather than compete with it. Thomas Wolf, who advises us at Bioptimus and helped build that playbook at Hugging Face, has always argued that open infrastructure compounds in ways closed products don't. I take that seriously, because I've watched it happen from the inside.

Why this doesn't come cheap

There's a reason the licensing gets stricter as the models get more powerful, and it isn't just business instinct — these models are extraordinarily expensive to build, in ways that get lost when people lump us in with "AI for biology" as a category. Some of that cost is R&D in the ordinary sense: we're not fine-tuning someone else's architecture on a new dataset, we're inventing the architectures, the training procedures, and the data-preprocessing methods needed to make tissue images, molecular data, and clinical outcomes learnable together in the first place. This requires an extraordinary density of talents, fluent in both AI research and biology. Some of it is compute, like everywhere else in this industry.

But the cost that actually sets us apart is data. AlphaFold and the wave of protein-generative and "virtual cell" models that followed it work from data that is, relatively speaking, abundant and public: a protein sequence, a single cell in a dish, preclinical measurements once removed from an actual patient. We're aiming a level higher and an order of magnitude harder — modeling real human patients at scale, with enough clinical depth that a model can start to predict how a specific patient will respond to a specific treatment. That's the question behind the roughly 90% of drugs that enter clinical trials and never make it out, and it's simply a different, harder problem than folding a protein. The data barely exists publicly. Spatial transcriptomics alone can run five figures per patient, and we need it paired with imaging and long-term outcomes across thousands of patients, not dozens. We license what we can from partners and generate the rest ourselves, and neither is cheap. That's the real reason M-Optimus stays commercial-only for now, not because we've grown precious about sharing, but because what makes it valuable to a drug developer is exactly what made it expensive to build, in a way a fine-tuned open-source checkpoint never is.

Where it actually costs us

I'd be lying to you, and to myself, if I stopped there. Openness has a near-term price, and I want to name it rather than gloss over it. First, to be fair to our own offer: a commercial license is often more than "the same weights, minus an Apache license." What you actually get is access to the people who built the model: our scientists sitting with your data to figure out why a result looks off, help integrating the model into a specific pipeline or assay, sometimes co-developing a variant tuned to a trial design we've never seen before. From my experience, that expertise is usually the difference between a model that runs and a model that changes a decision. And yet we still have real cases where a client evaluates that offer, likes it, and goes and uses the free version instead anyway — not because it's better for their use case, but because free is free, and the value of hands-on support and performance drop with a less advanced model is easy to underrate until you're missing it.. For a company like ours, that stings but doesn't threaten us. For an early-stage startup in this space, with no other revenue line and investors watching the runway, that same dynamic can be existential. I don't have a clean answer for that tension. Dual licensing is our attempt at a compromise — not a solution — and I expect we'll keep adjusting where we draw that line as we learn more about who actually needs the free tier versus who's just taking the path of least resistance. Keeping M-Optimus commercial-only, shared for now with a handful of partners rather than the whole field, is the same trade-off from the other direction: it means real, patient-relevant use cases go unexplored a little longer than my instincts would like, in exchange for actually understanding what the model is good for before we scale access to it.

Staged release is also a safety decision

Working at the frontier of AI and biology forces us to anticipate not only benefits but also risks, and releasing our models in stages is also a safety decision.

A model that reads a tumor's biology more deeply than most human experts is built to predict and diagnose — it doesn't design anything, and that distinction matters more than most of the current discourse on AI risks allows for. Still, when something is both new and genuinely more capable, like M-Optimus, I'd rather put it in front of a small number of partners we know well first — watching closely how it's used, and where it might be misused — before we widen the door. That's the same instinct that shaped H-Optimus-1's academic-only release, and it's why H-Optimus-0 only went fully open once we'd already learned, in public, how a model like it actually gets used in the wild. Capability and oversight should grow together, not in either order. It's a slower way to be open. I think it's the right one.

Why we're doing it anyway

I keep coming back to the same test: what gets our models into the hands of someone who can use it to help a patient, faster? Usually, the answer is to open the door wider, not narrower — and to be honest in public about the cases where that costs us something. A field that's serious about impact has to be able to hold both of those ideas at once. We're going to keep trying to.

Author
Author
Jean-Philippe Vert, PhD