Benchmarks
TideLM results
No TideLM training pass has completed, so there are no TideLM benchmark scores yet. Upstream SmolLM3 numbers are not copied into this table as if they were produced by this project.
| Checkpoint | Core diagnostic | Validation loss | Blind preference | Speed | Decision |
|---|---|---|---|---|---|
| Stable TideLM | — | — | — | — | none exists |
This page is generated from committed experiment and release metadata after a pass; dashes mean “not measured,” never zero.
Suite scope
The core suite contains 12 cases and is a regression smoke test, not a standard leaderboard. The CPU gauntlet uses a predeclared five-case subset during training to control cost. A promoted release should run all practical cases or state exactly which were omitted and why.
External evaluation roadmap
When compute allows, TideLM configs will add standard, version-pinned tasks such as IFEval, HumanEval-compatible execution in a hardened sandbox, GSM8K, ARC, TruthfulQA, and summarization quality review. Results will include harness commit, prompt template, shot count, generation parameters, sample count, and hardware. No roadmap item is a present result.