StartupHub.ai · Jun 2026
Long-form coverage of Cumulus Labs' Ion engine, GH200 optimization, benchmark claims, and the company's cloud-plus-software positioning.
Public company, workplace, funding, and market signals
Updated Jul 30, 2026
Cumulus Labs is a YC W26, San Francisco-based AI infrastructure startup building a unified production inference platform that combines routing, caching, observability, evaluation, fine-tuning, and custom hosting around its Ion inference engine.
Primary product
Cumulus — a unified inference platform / GPU cloud for production AI
Founded
2025
Headquarters
San Francisco, California, United States
Team size
1-5
Work style
Onsite
Industry
AI infrastructure
Sub-industry
GPU inference platform / serverless GPU cloud
Offices
0 jobs at Cumulus Labs
Check back later for new openings
Business model
Stage
seed
Total raised
$500K
Latest round
Convertible Note · Jan 2026
Latest amount
$500K
Jan 2026 · Y Combinator
Investors
Small, early, in-person team that emphasizes systems depth, fast shipping, and end-to-end ownership; job posts say they care more about how candidates think than the specific toolchain, and they use Claude Code heavily.
Work style
Onsite
Visa sponsorship
Limited
Compensation
Public hiring material for the founding ML platforms engineer listed $150K-$300K base pay, 1%-5% equity, and a $5K referral bonus; the role was full-time, in San Francisco, and restricted to US citizen/visa candidates.
Benefits
Pricing
Pay-per-million-tokens for popular supported models, with per-GPU-second pricing for custom, finetuned, or unsupported models; free tier available and volume pricing on request.
Differentiators
Technology
Customers
Competitors
StartupHub.ai · Jun 2026
Long-form coverage of Cumulus Labs' Ion engine, GH200 optimization, benchmark claims, and the company's cloud-plus-software positioning.
Agent Wars · Mar 2026
Independent coverage of Cumulus Labs' IonRouter launch, describing a GH200/B200-focused inference service, model multiplexing on one GPU, and the company's throughput and cold-start claims.
Cumulus Labs Blog · Feb 2026
Benchmarks and pricing comparison showing multiple vision-language models running on a single GPU, with per-token and per-GPU-second pricing positions and scale-to-zero economics.
Cumulus Labs Blog · Feb 2026
Technical deep-dive on Ionattention, the C++ inference runtime built for Grace Hopper, including custom CUDA-kernel techniques and benchmark claims such as 588 tok/s on a single GH200 in one cited comparison.
Y Combinator on LinkedIn · Jan 2026
Launch post describing Cumulus Labs as a YC W26 GPU cloud that optimizes training and inference workloads, charges by physical GPU usage, and highlights the founders' backgrounds.