GUIDES
Model keys
Agentifys does not resell tokens. You bring your own model key, and every call is billed by your provider, not by us.
Add a provider key
Go to Settings, API Keys, pick a provider, and paste your key. You can add more than one. The key is encrypted with Fernet before it is stored and is only decrypted in memory, at the moment a call is made. Agentifys supports three providers today.
| Field | Type | Description |
|---|---|---|
Anthropic | claude | Default model claude-haiku-4-5-20251001. |
OpenAI | gpt | Default model gpt-4o-mini. |
Google | gemini | Gemini through its OpenAI-compatible endpoint. Default model gemini-2.5-flash. |
The defaults are fast and cheap on purpose, a good fit for a support or in-product agent. Each key carries its own model, chosen from the dropdown next to it, so you can point one key at a larger model without touching the others. A key is usable the moment you save it: it starts at status unknown, which is not the same as failing, and it is tried on the next request.
Priority order and what actually fails over
Keys are tried in the order you set. Drag a key up to make it primary. Each request starts at the top of the list, skipping any key that is inactive or currently marked failing, and runs that key with that key's own model.
Fallback to the next key happens on authentication failure only. If the provider returns an authentication error, Agentifys marks that key failing, drops it from the list for this turn, and retries with the next key. Every other provider failure ends the turn on the key that was being used. Nothing is retried and no other key is tried.
| Field | Type | Description |
|---|---|---|
Authentication error | 401 / invalid key | Falls over to the next key. The failed key is marked failing with "Authentication failed — key may be revoked or expired". |
Rate limit | 429 | Ends the turn. No fallback. The user is told the provider is temporarily rate-limited. |
Quota / billing exhausted | provider error | Ends the turn. Providers report this as a rate-limit or generic error, not an authentication error, so it does not trigger fallback. |
Overload, timeout, bad request | other | Ends the turn with "Agent error:" and the provider message. |
When a turn ends on a rate limit, the stream carries this and then closes:
{
"type": "error",
"content": "The AI provider is temporarily rate-limited. Please wait a moment and try again."
}The daily health probe
A background loop verifies every active key across every project. It sleeps 24 hours after the backend starts, runs a pass, then repeats every 24 hours. Because the timer restarts with the process, a backend that is redeployed or restarted more often than once a day never completes a pass.
A pass sends each distinct key one real, minimal completion — the message hi with max_tokens of 1 — and records the outcome:
| Field | Type | Description |
|---|---|---|
Call succeeds | healthy | Status set to healthy. A key that was previously benched is put back into rotation by this. |
Call raises anything | failing | Status set to failing, with the first 200 characters of the provider error stored as last_error. This includes a rate limit, a timeout, or a transient provider outage, not only a bad key. |
claude-haiku-4-5-20251001, gpt-4o-mini, or gemini-2.5-flash — not the model you selected for that key. A key that works perfectly on its own model will be marked failing if its account cannot call the provider default, or if the provider happens to be rate-limiting or degraded during the pass.A key marked failing is excluded from every request until something clears it. There are exactly two ways to clear it:
| Field | Type | Description |
|---|---|---|
Test | manual | The flask button next to the key in Settings, API Keys. It calls that key with that key’s own model and sets the key back to healthy on success. This is the immediate fix. |
Next daily pass | automatic | The next probe re-tests failing keys too, and restores the key if the call succeeds. That is up to 24 hours away, and only if the process stays up that long. |
failing when the provider returns an authentication error; any other failure is reported back to you and the key's status is left alone.Alerts are in-app only
Per-project alerts are off until you enable them in the Analytics panel. Once enabled, a loop evaluates them every 5 minutes and fires at most one alert of each kind every 6 hours.
| Field | Type | Description |
|---|---|---|
key_failing | alert | Fires when any key on the project is marked failing. |
error_rate | alert | Errored turns as a share of turns over a rolling 24 hours, once there have been at least 5 turns. Threshold defaults to 25%. |
budget | alert | Month-to-date tokens against the monthly cap, when a budget is enabled. Threshold defaults to 90%. |
Why bring your own key
The tokens a conversation burns are billed to your provider account, at your negotiated rate, on your existing invoice. Agentifys never sits in the middle of that. There is no per-token markup and no surprise usage bill from us. You own the model relationship, the rate, and the data boundary.
It also means you choose the model. There is no shared platform key to fall back to: the server removes ANTHROPIC_API_KEY and OPENAI_API_KEY from its own environment at startup, so a project with no working key of its own cannot quietly spend someone else's.
When there is no working key
A project needs at least one usable key to answer. If it has no active key, or every active key is already marked failing, or every key authentication-fails during the turn, chat is rejected rather than silently degraded:
{
"type": "error",
"content": "PROJECT_NO_API_KEY"
}The fix is to add a working key, or to press Test on a benched key that you believe is fine. A key restored to healthy is picked up by the next request.