Models
Add a model service, choose a default model, and where keys are stored.
Rna connects three ways: OpenAI-compatible APIs (Chat Completions), Anthropic Messages, and signing in with a ChatGPT plan.
Settings include 17 provider presets that fill in the address and a verified model catalog:
| Group | Presets |
|---|---|
| China | Qwen (China / International), Kimi (China / International), Kimi Coding plan, Zhipu GLM API, Zhipu GLM Coding Plan, Xiaomi MiMo API, Xiaomi MiMo Token Plan, DeepSeek |
| International | OpenAI, Anthropic, Google Gemini, Groq, OpenRouter |
| Plan sign-in | OpenAI · ChatGPT plan |
| Other | Custom compatible service |
The Kimi Coding plan serves interactive chat only; background initiative and evolution need a separate pay-as-you-go provider.
In the UI
- Open Settings → Models (
⌘ ,). - Add a provider with its address and key. You can save a catalog with only those two first.
- Click “Test connection / fetch models”. It only lists models: no generation request, and discovered models are not added automatically; tick the ones you want and click “Add selected models”.
- Add models explicitly: model ID, name, context window, max output and supported reasoning levels. When “Fetch models after saving” is on (the default), saving a provider also adds the catalog models that were verified usable, up to 100. A model you removed is not added back automatically later; add it again when you need it.
- Make one the default, or switch per session.
A new turn pins the model configuration at that moment; a turn in progress keeps its own.
A successful connection test only proves the address and key work. It cannot prove the context size or reasoning levels you entered are valid. If the first real request fails, the error stays in the session trace; it is never dressed up as success.
From a file
To keep keys out of command-line arguments, write a provider file that references an environment variable:
{
"name": "My model service",
"protocol": "openai",
"baseUrl": "https://your-provider.example/v1",
"apiKeyEnv": "RNA_MODEL_API_KEY",
"models": [
{
"id": "your-model-id",
"name": "Project model",
"contextWindow": 128000,
"maxOutputTokens": 8192,
"reasoning": ["off", "high"]
}
]
}./rna provider-add --file /path/to/provider.json
./rna provider-test PROVIDER_ID
./rna model-default PROVIDER_ID MODEL_ID high
./rna model-set SESSION_ID PROVIDER_ID MODEL_ID highThe sizes and levels are format examples; fill in your service’s real values. The environment variable must be visible to the daemon. Setting it in another terminal does not change a daemon that is already running.
Where keys are stored
- Keys saved in the UI go to a local credentials file (mode
0600). Reading the configuration only returns “configured” and the source, never the key. - Keys never appear in normal state JSON, the event stream, browser localStorage or CLI output.
- The credentials file is still a plain local file; it is not in the macOS keychain yet.
- Remote addresses must use HTTPS; local addresses may use HTTP.
Custom OpenAI-compatible services
If your service supports JSON object output or tool_choice: none, declare it on the model:
"compat": { "supportsJsonObject": true, "supportsToolChoiceNone": true }Undeclared abilities are never switched on from the model name alone. false overrides automatic detection, and "compat": {} clears every explicit declaration.
With a ChatGPT plan
Choose OpenAI · ChatGPT plan, save it, then sign in through the browser or a device code. The model catalog comes from your account, and credentials refresh on each request.
- It uses your ChatGPT plan through the ChatGPT routes the Codex apps use. Those are not a public API and OpenAI may change them.
- The same sign-in can generate images and transcribe speech. The image models are gpt-image-2 (default), gpt-image-2.5-flare and gpt-image-2.5-sunburst, chosen under Settings → Models → Multimodal models, and the chosen model is what is actually sent; one image per call, up to five references. The account catalog lists no image models, so this list is built in. Speech synthesis and video have no such route and still need an API key.
- Generated images show directly in the conversation, the turn’s deliverables and the tool detail.
- Plan accounts and pay-as-you-go API keys are separate; switching never borrows the other credential.
With Claude
Claude connects only through an Anthropic API key, billed by usage. The preset includes Claude Opus 5.5, Sonnet 5 and Haiku 4.5.
- Claude conversation history is append-only, never rewritten, and thinking is bound to the conversation, so the prompt cache keeps hitting.
- While a conversation is idle, Rna re-sends the previous request just before the cache expires to keep it warm (no output is produced, and only when it is estimated to pay off); the spend counts against the real-token budget, and it stops on emergency stop, maintenance, deletion, a changed key or an exhausted budget. The usage page shows the keep-alive count under cache hits.
- Opus 5.5 and Sonnet 5 default to a 64K output allowance per request, with thinking counted in output; Opus 5.5 adaptive thinking cannot be switched off.
Anthropic allows Claude subscription logins only in Claude Code and Claude.ai. Third-party apps, including ones that drive the official Claude Code, are not permitted, and using a subscription there may get the account restricted. So Rna offers no Claude subscription sign-in, and the earlier read-only “Claude subscription consult” has been removed. The Claude subscription entry in settings explains why and switches to the API-key preset in one click.
Switching to a smaller context window
You can change models mid-conversation. When the new model’s window can’t hold the history, Rna compacts first:
- the summary reads a trimmed transcript: tool results, arguments and host context keep only their first part;
- when the service says a batch is too long, the batch is halved and retried, up to three times;
- when an answer request is refused as too long, it is compacted to half of what was sent and retried once.
Separately, when a conversation has been idle longer than its model's prompt cache survives and holds more than half of the context window, Rna compacts the earlier context into a summary in the background, keeping the last two turns, so the next turn does not wait for it. The original record and full tool results stay in the session log, and the conversation shows a note.
Reasoning effort of background and internal calls
Settings → Models → Reasoning effort of background and internal calls sets how deeply the calls Rna starts itself (background judgments, message rewriting and so on) think:
- Automatic (recommended): when the provider's default effort can adapt, or sits below the top level, it is used as is; when the default is the top level or unknown (GLM-5.3, for example), JEV picks one level per call, never above the conversation's own level.
- Same as the conversation: follow the reasoning level chosen for the conversation.
On providers that need JEV to pick, the choice is only recorded and does not take effect until a paid replay has passed. The conversation's own reasoning level also offers “Automatic” and “Provider default”.
Demo model
The built-in demo model needs no key and is only for trying the UI. Real models are called only after you switch to another provider.