AI chat with a local language model –
offline and private
Qwen, Gemma, Llama, GLM and others run directly on your machine. Your conversations are stored nowhere but with you, and nobody trains on them.
Alates 60 € · ühekordne makse, ilma tellimuseta · Windowsile, Linuxile ja macOS-ile
Põhiline
Mudelid ühe klõpsuga
A catalogue of around 50 variants; a filter hides what will not fit your video memory.
Choose per message
For every request you can decide which model should answer.
Thinking mode
With reasoning models you see the train of thought live before the answer arrives.
Even without a GPU
Very large models can run in system memory when they do not fit into video memory.
Tools without the cloud
From around 25 billion parameters a local model reaches for tools, web search and web reading reliably in the chat – the tool list in the prompt is deliberately kept compact for that.
Get models without hunting
What a local model actually delivers
Above says it runs offline. Here is what that means day to day – sizes, limits, and what works on which card.
Which models, and how large
Qwen, Gemma and Llama are the common families, plus around fifty variants in the library. For scale, taken from the program itself: qwen3:8b takes 5.2 GB, llama3.1:8b 4.9 GB, gemma3:12b 8.1 GB, qwen3:14b 9.0 GB. The filter hides what would not fit your card anyway and recommends the largest version that still runs comfortably.
Tools without the cloud – with a size threshold
Since v2.7.7 local models may use tools as well. From about 25 billion parameters a model gets the tool list in chat; below 30 billion a shortened core list of roughly 18 tools (3.4 KB instead of 52 KB), because smaller models otherwise start inventing calls. From 30 billion, and for cloud models, the full list applies.
How large the memory is – and why it varies
The context window is computed from two things: what the model can do, and what still fits on the card beside it. The smaller value wins. Next to a very large model, little is left over. If you want a large window, take a smaller model, a bigger card, or offload – a fixed high number would be a promise the hardware cannot keep.
When the history fills up
At 90 percent Qwirbel saves first instead of quietly truncating. Pinned lines survive the trimming. How many messages the short-term memory holds has been adjustable since v2.9.2 – the default is fifteen. More helps with "do another one", but costs tokens on every answer.
What offline really means
After the one-time activation you can pull the network cable and keep working. There is no service that has to answer for the program to start, no queue when others are busy, and no external provider's content filter refusing an answer to a harmless text.
When a cloud key is worth it
When your card is not enough for a task, you enter your own key and compute there – during API operation your card then stays completely free, even for side tasks like chat titles. The "Fully local" switch turns every provider back off with one click.
Mida selle kohta küsitakse
Which model should I pick?
That depends on your hardware. The model catalogue has a “fits my card” filter that hides anything too large for your video memory, so you are not guessing. Smaller models answer faster, larger ones more thoroughly.
Can local models use tools, or do I need an API key for that?
They can, and no key is required. In the chat a local model from around 25 billion parameters reaches for tools, web search and web reading reliably. We run Gemma 26B as a GGUF at a heavy quantisation (Q2) through the bundled Klecks engine – so it does not have to be the largest build of the model, which is why it also fits on a mid-range graphics card. In Work and Code, models below 30 billion get a compact core list of the most important tools; larger local models and cloud providers get the full one. The reason is simple: a very long tool list in the prompt makes smaller models stumble.
Do I need an account?
No. There is no sign-up and no mandatory account. The licence key is bound to the device once, after which Qwirbel keeps working offline.
Are my chats used for training?
No. Conversations stay on your machine. Nothing is transmitted to us and no training happens on your content.
Can I still use cloud models?
Yes, optionally: with your own API key you can add external providers. That is entirely your choice – it is not required, and image, video and music generation stay local either way.
Seda oskab Qwirbel ka
Kõik võimed elavad samas rakenduses – üks ost, üks võti.