The LLM Data Company · Jul 2026
Announces a 150-task equity-research benchmark across 143 companies and 10 sectors, with rubric-based grading and reference harnesses for evaluating agents end-to-end.
Public company, workplace, funding, and market signals
Updated Jul 30, 2026
The LLM Data Company is a 2025 San Francisco startup focused on post-training data, LLM evaluation, and RL environments for frontier AI teams; it appears to operate publicly under the product/brand name doteval and also uses the Paper Instruments label on parts of its site.
Primary product
doteval — an AI-assisted workspace for creating, versioning, and executing evals and RL rewards
Founded
2025
Headquarters
San Francisco, California, United States
Team size
1-10
Work style
Remote
Industry
Technology, Information and Internet
Sub-industry
Post-training data and LLM evaluation tooling
Offices
0 jobs at The LLM Data Company
Check back later for new openings
Business model
Stage
seed
Total raised
$500K
Latest round
Seed · Jun 2025
Latest amount
$500K
Jun 2025
Investors
Small, research-heavy, high-autonomy startup culture with fast iteration, direct lab exposure, and a strong focus on building production-grade evaluation and post-training infrastructure.
Work style
Remote
Visa sponsorship
Limited
Compensation
The YC posting for a Founding Research Engineer lists $150K-$200K base salary and 0.50%-1.00% equity; visa sponsorship is limited to U.S. citizens or existing visa holders.
Benefits
Differentiators
Technology
Customers
Competitors
Estimated revenue
≈$440K ARR (estimated, 2025)
The LLM Data Company · Jul 2026
Announces a 150-task equity-research benchmark across 143 companies and 10 sectors, with rubric-based grading and reference harnesses for evaluating agents end-to-end.
The LLM Data Company · Jun 2026
Compares rubric judges for non-verifiable RL tasks and argues for explicit binary criteria; reports strong results for full-rubric grading and Opus 4.7 on physician-labeled decisions.
The LLM Data Company · May 2026
Describes a post-trained Kimi K2.5 checkpoint that improves medical benchmark performance and notes tool-calling regressions without a harness.
The LLM Data Company · Mar 2026
Introduces the company’s first medical reasoning model, focused on clinical accuracy and reduced sycophancy; reports SOTA on HealthBench Hard.
The LLM Data Company · Dec 2025
Technical post on rollout/training mismatch in RL, with guidance on sampling settings and importance-sampling corrections.