Introducing the Bioptimus SDK: embeddings and expression from one pass over the slide
The Bioptimus SDK is available today:
Bioptimus builds foundation models that learn the dynamics of human biology, from cell to tissue to organ. The SDK is how research and ML teams put them to work: point it at a slide and a model, and it masks out background, tiles the slide, runs inference and writes structured results to disk.
For most labs, working with Bioptimus began with H-Optimus on Hugging Face. The family has since passed 1.8 million downloads and is cited in around 150 papers. The SDK is the next step for that community — the whole pipeline, from slide to result, built around the models you already use.
We are reducing time to value, from install to first inference
The models are the ones you already load. What changes is everything around them.
What you get back
Your own tissue mask, if you have one
If you already have a segmentation you trust, from a pathologist's annotation or your own tooling, the SDK runs from it instead of its own. The model sees exactly the tissue you chose.
Tile embeddings, and a slide-level embedding computed for you
Run a slide through H-Optimus and you get a vector per tile. The SDK then gives you one call to pool those tiles into a single slide-level embedding for a whole-slide vector, or to pool M-Optimus tile predictions into a bulk RNA-seq-like profile. That pooling is cheap post-processing after inference rather than something the model does — but it is post-processing you would otherwise write yourself, and have to write the same way every time for results to stay comparable.
A cohort is your study, in one place
In the SDK, a cohort is the single source of truth for a study: it pairs each slide with its patient, timepoints, labels, and any clinical metadata (and bulk RNA, when you have it). Build one from a folder of slides or a spreadsheet, and the SDK tracks per-slide processing status, so runs are fully reproducible and resume where they left off.
The code
A first slide, end to end (points the SDK at a running server and one slide):
To move from one slide to a cohort, and from embeddings to spatial expression, you change the arguments, not the code:
Under the hood
The Inference object handles reading, tissue masking, tiling, inference and writing. Tissue masks are cached, so repeat runs are faster and a run that stops part-way picks up where it left off.
Outputs are written as Zarr, HDF5 or NPZ, with tile coordinates, a thumbnail and the mask saved alongside. Every run's configuration is stored with its results, so a colleague — or a reviewer, eighteen months later — can reproduce exactly what you ran.
Go further with M-Optimus
The open weights give you embeddings. M-Optimus gives you gene expression: it predicts expression tile by tile, with or without bulk RNA, so you can see where in the tissue the signal sits. The SDK aggregates those tile predictions into a whole-slide pseudobulk profile.
It runs through the same SDK: change the model, keep your code. It is available through AWS Marketplace and on-premise — which for teams working under data-residency constraints means the slides never leave your environment.
Availability and licences
The SDK is available today with pip install bioptimus-sdk. Model licences are unchanged.
The SDK itself is released under a proprietary license. Bioptimus models are for research use and are not approved medical devices.