Skip to content
New NVIDIA B200 now available Reserve capacity
reviosa cloud
Engineering journal

The Reviosa Blog

GPU benchmarks, infrastructure deep dives, and product news from the team running the Reviosa fleet across four regions.

July 14, 2026
Product Featured

Introducing per-second billing for every instance type

Starting today, every Reviosa instance type bills by the second with a 60-second minimum. A nine-minute smoke test on an 8×H100 node now costs $2.75, not $18.32.

MR

Maya Reyes

CEO

Jul 14, 2026

5 min read

June 30, 2026
Guides

From notebook to production: vLLM endpoints on Reviosa

The distance between a working notebook and a production endpoint is smaller than it looks. Launch, serve, monitor, and load-test an OpenAI-compatible vLLM endpoint on Reviosa in an afternoon.

TB Tom Berger · 8 min read
June 9, 2026
Engineering

Why VRAM, not FLOPS, limits your batch size

Two H100 numbers matter more than TFLOPS ever will: 80 GB and 3.35 TB/s. A tour of where the bytes actually go in training and inference, and why memory hits the wall first.

SL Sofia Lindqvist · 7 min read
May 28, 2026
Product

On-demand vs reserved: choosing a GPU pricing model

Reserved capacity is meaningfully cheaper per hour, and on-demand is 100% cheaper when idle. The only question that matters is utilization — here's the break-even math.

MR Maya Reyes · 6 min read
May 12, 2026
Guides

Fine-tuning Llama 3 70B on a single 8×H100 node

You don't need a cluster to fine-tune a 70B model. A complete LoRA recipe for one 8×H100 node — memory budget, config, throughput, and the $64 bill at the end.

PS Priya Sharma · 10 min read
April 21, 2026
Engineering

Multi-node training over InfiniBand: a practical NCCL tuning guide

Multi-node throughput is won or lost in NCCL configuration before your training script runs a single step. The baseline-first workflow we use to get 90%+ scaling efficiency on InfiniBand clusters.

DO Daniel Okafor · 11 min read
March 31, 2026
Guides

Cutting inference costs 60%: quantization, batching, right-sizing

Most inference bills are 2-3× larger than they need to be. A practical walkthrough of quantization, continuous batching, and right-sizing that took a production Llama 3 8B deployment from $3,635 to under $1,300 a month.

TB Tom Berger · 9 min read
March 10, 2026
Benchmarks

H200 vs H100: real-world LLM training benchmarks

We ran identical FSDP fine-tuning jobs on 8×H100 and 8×H200 nodes and measured tokens per second, not marketing slides. Bandwidth buys more than you'd guess at long context — and less than the price delta everywhere else.

PS Priya Sharma · 8 min read
Get started

Start training in minutes

Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.

No minimum commitment · Cancel anytime · $10 free credit for new accounts