Sicey, end to end
One API key for every model. Sicey sizes each request by rules and sends it down your list of models: the hard ones stay on your best, the easy ones fall to something cheaper as its budget runs down.
Get started
- Open the app and make a workspace. You get a Sicey key that starts with
sk-sicey-. It is shown once and only a hash is kept, so copy it. - Put models in order on the Ladder page, coarsest first. A new workspace has an empty ladder on purpose: guessing three expensive models for you would route your first request somewhere you never chose.
- Add your own provider keys on the Providers page, or just an OpenRouter key, which reaches everything.
- Point your client at Sicey and send a request. The model you name is ignored.
Keys
A Sicey key is stored as a SHA-256 hash and compared in constant time. There is no route on this service that can return a key: a service able to show you your own key later is a service that can be made to show it to somebody else.
Provider keys are a different thing and are stored differently: encrypted at rest with AES-256-GCM, and only ever decrypted to make a call to that provider for you. If the server has no encryption secret configured it refuses to store a provider key at all rather than quietly keeping it in the clear.
You cannot revoke your last live key. That would lock you out of the workspace and the credit in it with no way back, so the route refuses and says so.
The API
Both wire formats work on both endpoints, so OpenAI and Anthropic clients need nothing special beyond a base URL.
# OpenAI clients OPENAI_BASE_URL=https://sicey.fun/v1 OPENAI_API_KEY=sk-sicey-... # Anthropic clients ANTHROPIC_BASE_URL=https://sicey.fun ANTHROPIC_API_KEY=sk-sicey-...
curl https://sicey.fun/v1/chat/completions \
-H "authorization: Bearer $SICEY_KEY" \
-H "content-type: application/json" \
-d '{"messages":[{"role":"user","content":"rename userId to accountId"}]}' -i| Endpoint | What it does |
|---|---|
| POST /v1/chat/completions | OpenAI chat completions, with tools. |
| POST /v1/messages | Anthropic messages, with tools. |
| POST /v1/messages/count_tokens | Exact when your top deck is an Anthropic model and you have an Anthropic key, because the real counter is asked. An estimate otherwise, and the reply says sicey_exact: false with how it was made. Nobody else publishes a counting endpoint, so there is no exact answer to give and pretending otherwise would be worse than saying so. |
| GET /v1/models | The models on your ladder, in order, with their prices. |
Streaming works on both endpoints, in both directions. Send stream: true and you get server-sent events in the format you asked in, even when the model that answers speaks the other one. The deck that answered is in the response headers before the first token arrives, so your client knows where it landed without waiting for the end.
One consequence worth knowing: a request can only fall to another deck before the first byte reaches you. Once text is flowing the connection is committed, so a model that dies mid-answer ends with an error frame rather than a silent truncation, because a truncated stream looks to every client like a complete short answer.
What comes back
The reply is in the format you asked in. These headers say what happened to it, and they are on every successful response.
| x-sicey-model | The model that actually answered. |
| x-sicey-rung | Its place on your ladder, 0 being the top. |
| x-sicey-size | What the request sized at, 0 to 1. |
| x-sicey-verdict | hard, medium, routine or easy. |
| x-sicey-because | One sentence saying why it landed there. |
| x-sicey-cost-cents | What that request cost, at the deck's own rate. |
| x-sicey-charged | credit, or your-own-key. |
| x-sicey-tried | Every deck tried, in order, when one refused. |
| x-sicey-dropped | Anything the format translation could not carry. Absent when nothing was. |
x-sicey-dropped is the one worth watching. Thinking blocks and cache markers belong to one model and cannot always cross between formats; when something is lost it is named rather than silently removed, because a dropped cache marker makes the next request cost ten times more with nothing in the output to show it.
The ladder
Your models, coarsest first, up to twelve. Sicey works down that list and nothing else: it will never quietly answer with a model your ladder does not name. If no key on the workspace can reach a deck, the request is refused with a sentence naming the missing key.
Sicey’s own GPU is added at the bottom automatically when it is serving, so the ladder always has a floor. A floor that cannot catch anything is not a floor, so it is left off when it cannot answer.
How sizing works
Three inputs, all local, all cheap:
- How much there is to do. Word count, with a ceiling at 40 words, so a pasted log file does not score as architecture.
- What kind of words. Two lists. Work that needs judgement (architect, migrate, refactor, deadlock, invariant, root cause) pushes up; work that needs typing (rename, typo, format, summarise, json, changelog) pulls down.
- Whether tools are attached. A request that can call tools is an agent step, and an agent step that picks the wrong tool is expensive to undo.
Only the last user message is sized. A long thread about a hard problem can end in “thanks, now rename it”, and sizing the whole transcript would send that to the most expensive model you own.
The number is compared against an aperture that opens as the top model’s daily budget drains. Above half the budget the aperture is shut and everything stays on top; below half it opens linearly, so the last of the budget goes only on the hardest work. The front page lets you drive both.
Credit and burning
The API works while the workspace has credit. Credit comes from burning the Sicey token from your own wallet: one transaction that burns the tokens and carries a memo naming your workspace. Sicey reads that transaction back from Solana before adding anything, so a burn credits only the workspace it names, and only once.
One million tokens adds $5.00 of credit.
That rate is a fixed number of cents. The token’s market price is not fixed. If a million tokens is worth more on the market than $5.00, burning it is a loss against selling it, and that can be a very large loss without anything on this page changing.
So the Credits page prints both numbers side by side, every time, read fresh: what a million is worth on the market, and what burning a million gives you. When the market price cannot be read it says so instead of showing our half on its own.
Sicey’s GPU
| Standard | Sicey GPU: Qwen3 30B A3B · $0.10 in, $0.40 out per million · 65,536 context |
| Pro | Sicey GPU Pro: gpt-oss-120b · $0.25 in, $1.00 out per million · 131,072 context |
These are our own prices on our own hardware, not a third party’s. They are not the cheapest way to reach an open-weight model, and we are not going to pretend otherwise: what they buy is a deck that needs no key of yours and catches work when everything else is spent.
Limits
- Twelve decks per ladder. More than that is a list nobody reads.
- Twenty live keys per workspace.
- A rate-limited deck is closed until the provider’s own
retry-aftersays, read from the header rather than guessed. - Sicey cannot see the limits of a ChatGPT, Claude.ai or Gemini subscription and does not change them. It watches the budget you set and the errors your providers return.