Agent traces you already capture are opportunities to get signal on how to make your model cheaper, faster, better.
We do this by continuously a) distilling relevant chain of thought from larger open source models into smaller ones, b) model routing to frontier + OS models, and c) token compaction to remove noise and save on tokens.
`wmo build` allows you to build a simulation to optimize against with your agent traces with your OpenRouter key
`wmo optimize` trains a router, compaction, and distills chain of thought from a larger model into your specialized model
`wmo serve` gives you an endpoint for your model
When you call your model, behind the scenes a router decides which tasks should go to the frontier versus your model. Tinker continually trains as new traces arrive.
We're also working on a hosted solution that does continual training + serving for you https://experientiallabs.ai
Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about.
Waitlists are against the Show HN rules (https://news.ycombinator.com/showhn.html), and you're likely to get a lot of community pushback if you post before there's enough substance for users to sink their teeth into.
wmo routes requests between frontier models and open source models that continuously train using Tinker. As the smaller models improve, more traffic gets routed to them.
Calculating cost is just tokens in/out.
We do have a platform we'll be launching as well to manage training + serving for you which will require more diligent privacy guarantees.
The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes.
Until then, this is noise.