Quill & Query serves contract analysis at sub-300ms p95 — and pays nothing when lawyers sleep
“Legal traffic is the spikiest workload I've ever run. Nothing at 2 a.m., then an entire M&A data room lands at 9:05 on a Monday. Serverless endpoints on Reviosa gave us hardware that matches that reality — instant when we need it, free when we don't.”
Dana Okafor
CTO, Quill & Query
Quill & Query builds AI-assisted contract analysis for mid-size law firms. Attorneys upload agreements — often hundreds at once during due diligence — and the platform extracts clauses, flags deviations from firm playbooks, and answers natural-language questions with citations back to the source text. Under the hood is a retrieval-augmented generation pipeline: an embedding model for indexing, a reranker, and a fine-tuned language model that drafts grounded answers.
The challenge
Law firm traffic has a shape that punishes always-on infrastructure. Usage is concentrated in business hours, spikes violently when a deal's document set arrives, and drops to nearly zero on nights and weekends. Quill & Query's original deployment — a fixed pool of GPU instances sized for peak — sat below 15% utilization on an average day. More than two-thirds of the serving budget bought idle silicon.
Shrinking the pool wasn't an option either. Attorneys billing by the hour do not wait for cold starts; the product's contract commits to responsive interactive Q&A, and internal targets called for sub-300ms p95 on retrieval-augmented answers' first token. The team needed peak-grade latency and valley-grade economics from the same infrastructure.
Why Reviosa
Reviosa's serverless inference endpoints resolved the contradiction. Quill & Query deployed its embedding, reranking, and generation models as three serverless endpoints that scale with request volume — including all the way to zero. Warm capacity is managed by the platform, so the Monday-morning surge lands on ready GPUs instead of cold ones, and per-second billing means a quiet Sunday genuinely costs nothing.
The heavier lifting stayed on dedicated capacity where it belongs: nightly index rebuilds and quarterly model fine-tunes run on on-demand H100 instances at $2.49/hr, spun up by the batch scheduler and terminated on completion. One platform, two very different workload shapes, each billed the way it actually behaves.
"We stopped writing capacity-planning documents. The endpoints follow the traffic, and the invoice follows the endpoints. Our infrastructure meetings got very short." — Dana Okafor, CTO
Migration took three weeks, most of it spent on load testing rather than re-engineering — the models deployed from the same containers the team already built for its fixed pool.
The results
Six months in, the numbers have held through two of the busiest deal seasons in the company's history:
- 287ms p95 latency on interactive contract Q&A, comfortably inside the 300ms target — measured during peak due-diligence surges, not quiet periods.
- $0 idle cost. Endpoints scale to zero outside business hours; overnight and weekend serving spend fell to effectively nothing.
- 70% lower serving spend overall compared with the fixed GPU pool, even as monthly query volume grew 3x.
Reliability improved alongside the economics. When a single firm uploaded an 11,000-document data room — 40x normal load — the endpoints absorbed the spike without paging anyone. Under the old fixed pool, that event would have meant emergency capacity and degraded latency for every other customer.
The savings went straight back into the product. Quill & Query used the reclaimed budget to fine-tune a larger generation model — trained on Reviosa H100s — that lifted clause-extraction accuracy enough to expand into two new practice areas. For a 26-person company selling to risk-averse law firms, that is the whole story: the infrastructure finally costs what the traffic costs, and the difference funds the roadmap.
The results, by the numbers
287ms
p95 response latency
$0
idle infrastructure cost
70%
lower serving spend
More teams on Reviosa
Loomline AI cut diffusion training costs 55% without changing a line of code
Generative AI / Fashion55%
lower training cost
Helixon Bio compressed a 6-month screening campaign into 3 weeks on burst H200 capacity
Biotech / Drug Discovery2M
compounds screened
Northfork Render finished a feature film on 400 on-demand L40S GPUs — with zero hardware owned
Media & Entertainment / VFX400
L40S GPUs at peak
Start training in minutes
Create an account, add a card, and launch your first GPU instance. Per-second billing means you only pay for what you use.
No minimum commitment · Cancel anytime · $10 free credit for new accounts