
Sicey
Hard work stays on the top deck.
A sizing line is a stack of screens, coarsest on top. Everything is tipped onto the first one; what is too big to pass stays there, and the rest falls through until a screen holds it.
Your models are the screens. One API key for all of them, and the hard requests keep your best one while the easy ones go somewhere cheaper as its daily budget runs down.
Your best model keeps the hard work
Sicey sizes every request before it goes out and sends it down your list of models: the hardest stay on top, the easy ones fall to cheaper decks as the top one’s budget runs low. By rules, not by asking another model to read your prompt first.

Your best model
Plans, hard bugs, long refactors. Sicey keeps these here until the daily budget you set is nearly gone.
Claude Fable 5.1, GPT-6 Astra

A mid-size model
Most coding work and most tool calls, once the top model's budget is past half.
Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro

A small, fast model
Renames, summaries, lookups, format changes: work that needs typing rather than judgement.
Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4.1 Flash

Sicey's own GPU
The floor. No provider key needed, and it catches everything when the rest are spent or rate-limited.
Paid from credit

A model of its own
Sicey runs its own models on its own hardware. They answer with no provider key at all, so when every model on your ladder is spent or rate-limited the work goes there instead of stopping. Every reply names the model that wrote it, in a header, on every request.
Access is paid for by burning the Sicey token from your wallet, which turns into credit for the API.
Not serving on this deployment. SICEY_GPU_URL is not set, so there is no worker to send a request to.
Not serving on this deployment. SICEY_GPU_URL is not set, so there is no worker to send a request to.
Setup is one setting
Make a key in the app, point your client’s base URL at Sicey, and send any model name you like: the one you send is ignored, because picking it is the job. Add your own provider keys whenever you want Sicey routing between your own models.
Which models does it work with?
Anything you can reach with an Anthropic, OpenAI, Google or OpenRouter key. OpenRouter covers the rest: Grok, DeepSeek, Mistral, Qwen, Llama, Kimi, GLM and hundreds more. You put the models in order on the Ladder page and Sicey works down that list, and only that list.
How does it decide how hard a request is?
By rules, not by asking another model to read your prompt. Three inputs: how long the last message is, with a ceiling so a pasted log file is not mistaken for architecture; which words are in it, from two lists, one of work that needs judgement and one of work that needs typing; and whether tools are attached. The result is one number between 0 and 1, and the dashboard shows the arithmetic for every request.
How do I pay for Sicey?
By burning the Sicey token on Solana. Burning it from your wallet on the Credits page adds credit to your workspace, and the API only works while there is credit. One million tokens adds $5.00. Sicey's own GPU draws on that credit; routing to your own provider keys does not use it up, because that request is billed to you by the provider already and charging for it twice would be charging twice.
Does burning cost more than the credit is worth?
Sometimes, and the Credits page tells you which. The rate is a fixed number of cents; the token's market price moves. So the page prints what a million tokens is worth on the market right next to what burning a million gives you, and when the market read fails it says the price is unknown rather than showing our half of the comparison on its own.
Which limits does it watch?
The daily budget you set for your top model, counted from the cost of the requests Sicey itself routed, and rate-limit errors. When a model returns a 429, Sicey reads the provider's own retry-after header rather than guessing a backoff, closes that deck until then, and the request falls to the next one. It does not change the limits of a ChatGPT, Claude.ai or Gemini subscription, and it cannot see them.
Does a conversation stay on one model?
Yes. A conversation stays on the model it started on and only moves when that model is closed, because thinking blocks and prompt caches belong to one model. Throwing a warm cache away to save a tenth of a cent makes the request more expensive, not less.
What happens to my keys?
Provider keys are encrypted at rest with AES-256-GCM and are only ever used to call that provider for you. There is no route that can hand one back. Sicey records the model, the token counts, the cost and the timing of each request: there is no column in the database that could hold a prompt or a reply.
Why not just lower the effort setting?
Try that first. A lower reasoning effort on one model often costs less and keeps a single prompt cache, which is two advantages Sicey cannot give you. Sicey is for when that still runs out before the day does.
- Price list
- 467 models from OpenRouter’s public list, plus 2 of our own.
- Read at
- 2026-10-08T10:06:34.196Z · under an hour ago.
That timestamp is written by the fetch itself, not typed into this page. If the refresh stops running, the date goes stale in public, which is the only kind of stale that gets fixed. - The artwork
- Generated for this site with npm run art, not photographed and not borrowed. Every machine stands on a tiered steel deck because that is what this product is named after.
- Domain
- https://sicey.fun