Osmosis blog · Jun 2026
Describes a fused Triton logprob kernel that reduced RL training time and increased trainable context length for long-context reinforcement learning.
Public company, workplace, funding, and market signals
Updated Jul 30, 2026
San Francisco-based YC W25 startup founded in 2024 by Kasey Zhang and Baiqing Lyu, building a forward-deployed reinforcement fine-tuning platform for AI agents and task-specific models that improve from real usage.
Primary product
Forward-deployed reinforcement fine-tuning / post-training platform for AI agents
Founded
2024
Headquarters
San Francisco, California, United States
Team size
1-10
Industry
Technology, Information and Internet
Sub-industry
AI infrastructure / reinforcement learning platform for agentic AI
Offices
11 jobs at Osmosis
Business model
Investors
Technical, customer-facing, and hands-on; the company emphasizes bleeding-edge work, direct support with customers, high agency, and a paid in-person work trial in hiring.
Visa sponsorship
Limited
Compensation
Public hiring signals show full-time ML roles at $180K-$250K base salary in San Francisco; no public equity or bonus details were found.
Pricing
Not publicly disclosed; appears to be sold as a B2B platform / enterprise software offering.
Differentiators
Technology
Customers
Competitors
Osmosis blog · Jun 2026
Describes a fused Triton logprob kernel that reduced RL training time and increased trainable context length for long-context reinforcement learning.
Osmosis blog · Jun 2026
Explains a multi-LoRA/Megatron-Bridge approach that enabled thousands of concurrent LoRA adapter trainings and improved throughput.
Signalbase · Oct 2025
Reports the company’s $6.3M seed round and says the capital will be used to scale operations, expand engineering and research, and accelerate product R&D.
Osmosis blog · Jul 2025
Introduces Osmosis-Apply-1.7B, an RL-tuned model for code merging that the company says outperformed larger foundation models on its validation benchmark.
Y Combinator
YC launch page describing Osmosis as infrastructure that lets AI agents learn from tasks in real time, with claims about accuracy, cost-efficiency, and fewer steps per task.