Helixon Bio compressed a 6-month screening campaign into 3 weeks on burst H200 capacity
“Our chemists used to design around the compute queue — we'd trim a screening library because the schedule couldn't absorb it. Running two million compounds in three weeks changed what questions we're allowed to ask. That's not an infrastructure win, that's a science win.”
Dr. Priya Raghavan
Head of Computational Chemistry, Helixon Bio
Helixon Bio is a preclinical drug-discovery company targeting fibrotic disease. Its discovery engine pairs structure-based virtual screening with ML-guided binding-affinity prediction: a transformer-based scoring model ranks candidate molecules, and the most promising hits flow into physics-based free-energy calculations before anything touches a wet lab. The bottleneck was never the science — it was getting enough memory-rich GPUs at the moment a campaign kicked off.
The challenge
Helixon's affinity models are large and long-context: a single scoring pass holds a protein pocket representation, a conformer ensemble, and a sizeable KV cache in memory at once. On the 40 GB and 80 GB cards in their on-prem cluster, the team had to shard models and batch conservatively, which capped throughput at roughly 80,000 compounds per week. At that rate, the flagship campaign — a 2-million-compound library against three novel targets — was scheduled to take six months. In drug discovery, six months of compute time is six months a competitor can use to reach the same target first.
Buying more hardware was the obvious answer and the wrong one. Campaigns are bursty: Helixon needed enormous capacity for a few weeks, then almost none while chemists validated hits.
Why Reviosa
Reviosa's H200 SXM instances were the fit the team had been sketching on whiteboards. With 141 GB of HBM3e per GPU at $3.69/hr, an entire scoring model plus its working set fits on a single card — no sharding, no cross-GPU communication overhead, and batch sizes four times larger than the on-prem setup allowed. Because Reviosa bills by the hour with no minimum commitment, Helixon could treat the fleet like a lab instrument: reserve nothing, burst hard, release everything.
The team wrapped its screening pipeline in a simple queue that scaled Reviosa instances through the API. Compound batches streamed from object storage; results landed back as Parquet files the chemists could query the same afternoon.
"We stopped thinking in quarters and started thinking in weeks. Capacity that used to be a capital-planning meeting is now an API call." — Dr. Priya Raghavan, Head of Computational Chemistry
The results
The flagship campaign ran end to end in 21 days:
- 2 million compounds screened across three targets, with the full ML scoring pass and top-decile rescoring completed in a single burst window.
- 96 H200 GPUs at peak, scaled up over two days and released completely when the campaign closed — the cluster cost Helixon nothing the week after.
- 9x throughput per GPU versus the on-prem 80 GB cards, driven by unsharded models and larger batches on 141 GB of memory.
The compressed timeline mattered beyond the calendar. Because results arrived while the underlying biology was still fresh, Helixon's chemists iterated twice more on the library design within the same quarter — something the six-month plan made impossible. Two of the three targets produced advanced hit series now moving into lead optimization.
Helixon has since made burst screening its standard operating model. Every new target program budgets a three-week Reviosa window instead of a six-month cluster reservation, and the on-prem hardware now handles what it is actually good at: steady, predictable free-energy refinement between campaigns. Total compute spend for the flagship campaign came in 38% below the projected cost of running it in-house — before counting the five months of calendar time the company got back.
The results, by the numbers
2M
compounds screened
3 wks
vs. 6-month baseline
96
H200 GPUs at peak
More teams on Reviosa
PixelPatch scaled AI photo enhancement to 4M requests a day while cutting infra spend 40%
Consumer Apps4M
requests per day
Loomline AI cut diffusion training costs 55% without changing a line of code
Generative AI / Fashion55%
lower training cost
Northfork Render finished a feature film on 400 on-demand L40S GPUs — with zero hardware owned
Media & Entertainment / VFX400
L40S GPUs at peak
Start training in minutes
Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.
No minimum commitment · Cancel anytime · $10 free credit for new accounts