Releases: FRENCHIIIFRIES/emb3r-ai
Release list
emb3r v1.38.1
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.38.1 — she tells you when she cannot speak, and gives the memory back
When Ember could not speak, she said nothing at all. No sound and no reason, which looks exactly like the feature not existing — and is why the only thing anyone could report was "I can't hear her". There were three separate ways for this to happen silently and all three now say what went wrong. On a machine short of memory the answer is usually that there was not enough free to load her voice.
Replies should be quicker again. Her voice and her hearing were staying loaded for as long as emb3r was open. Measured, the voice alone takes the app from 59 MB to 343 MB — memory the model then has to think without, which on a laptop with half a gigabyte free is the difference between thinking and waiting on the disk. Both are now let go after a minute and a half unused, which gives back 130 MB of it, and neither loads until it is actually needed rather than on the chance.
If replies are still slow, the thing that will help most is switching to Qwen2.5 0.5B in Settings → Models. At 0.4 GB it is the only model in the list that comfortably fits a machine with very little free.
One thing that was tried and not shipped: pinning how many processor threads the model uses. Measured across five settings, the default it already chooses was fastest, and every value tested made it slower. Left alone.
v1.38.0 — she talks back in Talk, and two models that fit anywhere
Ember was silent in Talk. That view exists to be listened to, and she said nothing in it unless you had found a switch in Settings that is off by default. She always speaks there now. The setting still decides whether replies are read aloud in the terminal — and it now says that is what it does, rather than implying it governs everything.
You can speak to her in the terminal as well. Hold the [o] button beside the message box, say something, let go. What she heard goes into the box so you can read it before you send it — in Talk it is sent straight away, because there is no keyboard there to fix a misheard word with.
Two much smaller models. The smallest thing in the list was 1.9 GB, which is no use on a machine that has half a gigabyte free. Qwen2.5 0.5B is 0.4 GB and answered a short question in about seven seconds on exactly such a machine. Gemma 3 1B is 0.8 GB, takes more care over its wording, and takes longer to do it.
Both were asked real questions before being added, on a laptop in that state, and the times above are what they actually did rather than anything quoted from a description.
A third was tested and left out. Qwen3 1.7B is a year newer than anything else here and it loads perfectly — and asked three ordinary questions it produced nothing at all, three times over, spending about a minute each time thinking privately. Newer is not better on its own.
Neither of the new models knows anything about the world after it was trained, and no model that runs on your machine does. That is what web access is for.
v1.37.0 — your provider answers when you want it to
If you had set up Groq or another provider, it was answering everything — including "hello". The model on your own machine had stopped being used at all, which is the wrong way round for an application whose whole point is that it runs locally.
That was an overcorrection to the opposite problem a few versions ago, when a provider you had configured was never reached and appeared to do nothing. Rather than swing back and get it wrong in the other direction, Web access now asks.
Under the provider fields there are two choices. Only questions that need current information is the default: Ember answers everything on your machine, and your provider handles the questions she cannot know the answer to — the same rule Gemini has always followed. Every message sends everything to your provider, which is worth having if your machine is slow, and it says plainly that nothing you type stays local.
If you had a provider set up before this update you will be moved to the first option, which is most likely the change you wanted. If you preferred it the other way, the setting is two clicks.
Nothing about Gemini changed, and neither did the offline lock, the consent prompt, or the indicator that tells you when something is leaving your machine.
v1.36.0 — talking works in the installed app, and Macs can update again
Two of these are apologies.
Talking to Ember never worked in the installed app. It worked here while it was being built, and it worked in every test — but the version you download was missing one piece the speech recogniser loads before it does anything, so holding the button and speaking always ended in an error about a missing package. It has been broken since voice shipped, on Windows and on Mac. It is included now, and this time the check was done on a real installed copy rather than on the development one.
Updates install themselves on a Mac again. macOS refuses to let an application replace itself unless it was signed with an Apple certificate, which emb3r does not have — so the built-in updater was being turned away at the last step, after downloading the whole thing. emb3r now does that part itself: it downloads the new version, checks it against the checksum published alongside the release, and puts it in place. If anything fails at the final step the old version is put back rather than left half-replaced.
Worth knowing what that trade is. The Apple certificate is a promise that the app came from who it says; without one, the check is that the download matches the checksum the release published over an encrypted connection. That is a real check, and it is not the same check.
History, Talk and Settings are one menu now. They were three buttons crowding the corner with the network light — the one thing up there that is a promise rather than a control. The light has the room now.
v1.35.0 — a bigger face, and replies that are not waiting on the disk
Two things: Ember's face got much larger and learned five new expressions, and a memory bug that made replies crawl on a busy machine is fixed.
Press Face in the top bar and she fills the window. She has always had a set of expressions tied to what the app is actually doing — you have only ever seen them at the size of a line of text. Now they are the size of the window.
Five new ones, all about talking to her. Wide-eyed while the microphone is open, thoughtful while she works out what you said, puzzled when she could not make it out. If your microphone is missing or refused she tells you with her face rather than with an error. Before this, "listening" borrowed the surprised expression — the one view built for not reading the screen was showing you the wrong thing.
She takes your colour. The face is painted with a gradient built from whatever accent you have picked, lighter at the top and deeper at the bottom. If you have not picked one, she burns in emb3r's own fire instead of the interface green.
Replies could take minutes on a machine that was low on memory, and that is fixed. emb3r decides how much conversation to hold based on how much memory it can spare. It was working that out from how much memory your computer has in total, which assumes emb3r is the only thing running. On a laptop with everything else open it was asking for four and a half gigabytes when well under one was actually free, and Windows made up the difference by shuffling to disk. That is not slow thinking, it is a machine out of room — and it looks identical from the outside.
It now looks at what is genuinely free as well, and takes less when the machine is busy. If you have memory to spare nothing changes at all; the smaller window is only chosen when the alternative is swapping.
What went out. Face mode briefly had a drawn creature in it, with its own animation loop. Three designs were tried and none earned their place — one managed to be unsettling, another childish. That is a lot of machinery to maintain for something the faces were already doing well, so it is gone.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came ...
emb3r v1.38.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.38.0 — she talks back in Talk, and two models that fit anywhere
Ember was silent in Talk. That view exists to be listened to, and she said nothing in it unless you had found a switch in Settings that is off by default. She always speaks there now. The setting still decides whether replies are read aloud in the terminal — and it now says that is what it does, rather than implying it governs everything.
You can speak to her in the terminal as well. Hold the [o] button beside the message box, say something, let go. What she heard goes into the box so you can read it before you send it — in Talk it is sent straight away, because there is no keyboard there to fix a misheard word with.
Two much smaller models. The smallest thing in the list was 1.9 GB, which is no use on a machine that has half a gigabyte free. Qwen2.5 0.5B is 0.4 GB and answered a short question in about seven seconds on exactly such a machine. Gemma 3 1B is 0.8 GB, takes more care over its wording, and takes longer to do it.
Both were asked real questions before being added, on a laptop in that state, and the times above are what they actually did rather than anything quoted from a description.
A third was tested and left out. Qwen3 1.7B is a year newer than anything else here and it loads perfectly — and asked three ordinary questions it produced nothing at all, three times over, spending about a minute each time thinking privately. Newer is not better on its own.
Neither of the new models knows anything about the world after it was trained, and no model that runs on your machine does. That is what web access is for.
v1.37.0 — your provider answers when you want it to
If you had set up Groq or another provider, it was answering everything — including "hello". The model on your own machine had stopped being used at all, which is the wrong way round for an application whose whole point is that it runs locally.
That was an overcorrection to the opposite problem a few versions ago, when a provider you had configured was never reached and appeared to do nothing. Rather than swing back and get it wrong in the other direction, Web access now asks.
Under the provider fields there are two choices. Only questions that need current information is the default: Ember answers everything on your machine, and your provider handles the questions she cannot know the answer to — the same rule Gemini has always followed. Every message sends everything to your provider, which is worth having if your machine is slow, and it says plainly that nothing you type stays local.
If you had a provider set up before this update you will be moved to the first option, which is most likely the change you wanted. If you preferred it the other way, the setting is two clicks.
Nothing about Gemini changed, and neither did the offline lock, the consent prompt, or the indicator that tells you when something is leaving your machine.
v1.36.0 — talking works in the installed app, and Macs can update again
Two of these are apologies.
Talking to Ember never worked in the installed app. It worked here while it was being built, and it worked in every test — but the version you download was missing one piece the speech recogniser loads before it does anything, so holding the button and speaking always ended in an error about a missing package. It has been broken since voice shipped, on Windows and on Mac. It is included now, and this time the check was done on a real installed copy rather than on the development one.
Updates install themselves on a Mac again. macOS refuses to let an application replace itself unless it was signed with an Apple certificate, which emb3r does not have — so the built-in updater was being turned away at the last step, after downloading the whole thing. emb3r now does that part itself: it downloads the new version, checks it against the checksum published alongside the release, and puts it in place. If anything fails at the final step the old version is put back rather than left half-replaced.
Worth knowing what that trade is. The Apple certificate is a promise that the app came from who it says; without one, the check is that the download matches the checksum the release published over an encrypted connection. That is a real check, and it is not the same check.
History, Talk and Settings are one menu now. They were three buttons crowding the corner with the network light — the one thing up there that is a promise rather than a control. The light has the room now.
v1.35.0 — a bigger face, and replies that are not waiting on the disk
Two things: Ember's face got much larger and learned five new expressions, and a memory bug that made replies crawl on a busy machine is fixed.
Press Face in the top bar and she fills the window. She has always had a set of expressions tied to what the app is actually doing — you have only ever seen them at the size of a line of text. Now they are the size of the window.
Five new ones, all about talking to her. Wide-eyed while the microphone is open, thoughtful while she works out what you said, puzzled when she could not make it out. If your microphone is missing or refused she tells you with her face rather than with an error. Before this, "listening" borrowed the surprised expression — the one view built for not reading the screen was showing you the wrong thing.
She takes your colour. The face is painted with a gradient built from whatever accent you have picked, lighter at the top and deeper at the bottom. If you have not picked one, she burns in emb3r's own fire instead of the interface green.
Replies could take minutes on a machine that was low on memory, and that is fixed. emb3r decides how much conversation to hold based on how much memory it can spare. It was working that out from how much memory your computer has in total, which assumes emb3r is the only thing running. On a laptop with everything else open it was asking for four and a half gigabytes when well under one was actually free, and Windows made up the difference by shuffling to disk. That is not slow thinking, it is a machine out of room — and it looks identical from the outside.
It now looks at what is genuinely free as well, and takes less when the machine is busy. If you have memory to spare nothing changes at all; the smaller window is only chosen when the alternative is swapping.
What went out. Face mode briefly had a drawn creature in it, with its own animation loop. Three designs were tried and none earned their place — one managed to be unsettling, another childish. That is a lot of machinery to maintain for something the faces were already doing well, so it is gone.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came out the same length to within eight milliseconds.
A word about waiting. Making speech is slower than saying it, so Ember banks a couple of seconds of audio before she starts. Ordinary replies then run without a break. A very long answer can still catch up with her and pause part-way through; that is arithmetic rather than a fault, and it is better than what it replaced, which was a five-second stall in the middle of a sentence.
The download is bigger — about 135 MB. That is the two speech models, and they are in there so that all of the above works the first time you open emb3r, with the offline lock switched on, having fetched nothing at all.
The microphone is only ever open while you hold the button, there is an indicator beside the network one showing when it is live, and emb3r now turns down camera, location and notification requests by name rather than leaving them to a default it never chose.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card t...
emb3r v1.37.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.37.0 — your provider answers when you want it to
If you had set up Groq or another provider, it was answering everything — including "hello". The model on your own machine had stopped being used at all, which is the wrong way round for an application whose whole point is that it runs locally.
That was an overcorrection to the opposite problem a few versions ago, when a provider you had configured was never reached and appeared to do nothing. Rather than swing back and get it wrong in the other direction, Web access now asks.
Under the provider fields there are two choices. Only questions that need current information is the default: Ember answers everything on your machine, and your provider handles the questions she cannot know the answer to — the same rule Gemini has always followed. Every message sends everything to your provider, which is worth having if your machine is slow, and it says plainly that nothing you type stays local.
If you had a provider set up before this update you will be moved to the first option, which is most likely the change you wanted. If you preferred it the other way, the setting is two clicks.
Nothing about Gemini changed, and neither did the offline lock, the consent prompt, or the indicator that tells you when something is leaving your machine.
v1.36.0 — talking works in the installed app, and Macs can update again
Two of these are apologies.
Talking to Ember never worked in the installed app. It worked here while it was being built, and it worked in every test — but the version you download was missing one piece the speech recogniser loads before it does anything, so holding the button and speaking always ended in an error about a missing package. It has been broken since voice shipped, on Windows and on Mac. It is included now, and this time the check was done on a real installed copy rather than on the development one.
Updates install themselves on a Mac again. macOS refuses to let an application replace itself unless it was signed with an Apple certificate, which emb3r does not have — so the built-in updater was being turned away at the last step, after downloading the whole thing. emb3r now does that part itself: it downloads the new version, checks it against the checksum published alongside the release, and puts it in place. If anything fails at the final step the old version is put back rather than left half-replaced.
Worth knowing what that trade is. The Apple certificate is a promise that the app came from who it says; without one, the check is that the download matches the checksum the release published over an encrypted connection. That is a real check, and it is not the same check.
History, Talk and Settings are one menu now. They were three buttons crowding the corner with the network light — the one thing up there that is a promise rather than a control. The light has the room now.
v1.35.0 — a bigger face, and replies that are not waiting on the disk
Two things: Ember's face got much larger and learned five new expressions, and a memory bug that made replies crawl on a busy machine is fixed.
Press Face in the top bar and she fills the window. She has always had a set of expressions tied to what the app is actually doing — you have only ever seen them at the size of a line of text. Now they are the size of the window.
Five new ones, all about talking to her. Wide-eyed while the microphone is open, thoughtful while she works out what you said, puzzled when she could not make it out. If your microphone is missing or refused she tells you with her face rather than with an error. Before this, "listening" borrowed the surprised expression — the one view built for not reading the screen was showing you the wrong thing.
She takes your colour. The face is painted with a gradient built from whatever accent you have picked, lighter at the top and deeper at the bottom. If you have not picked one, she burns in emb3r's own fire instead of the interface green.
Replies could take minutes on a machine that was low on memory, and that is fixed. emb3r decides how much conversation to hold based on how much memory it can spare. It was working that out from how much memory your computer has in total, which assumes emb3r is the only thing running. On a laptop with everything else open it was asking for four and a half gigabytes when well under one was actually free, and Windows made up the difference by shuffling to disk. That is not slow thinking, it is a machine out of room — and it looks identical from the outside.
It now looks at what is genuinely free as well, and takes less when the machine is busy. If you have memory to spare nothing changes at all; the smaller window is only chosen when the alternative is swapping.
What went out. Face mode briefly had a drawn creature in it, with its own animation loop. Three designs were tried and none earned their place — one managed to be unsettling, another childish. That is a lot of machinery to maintain for something the faces were already doing well, so it is gone.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came out the same length to within eight milliseconds.
A word about waiting. Making speech is slower than saying it, so Ember banks a couple of seconds of audio before she starts. Ordinary replies then run without a break. A very long answer can still catch up with her and pause part-way through; that is arithmetic rather than a fault, and it is better than what it replaced, which was a five-second stall in the middle of a sentence.
The download is bigger — about 135 MB. That is the two speech models, and they are in there so that all of the above works the first time you open emb3r, with the offline lock switched on, having fetched nothing at all.
The microphone is only ever open while you hold the button, there is an indicator beside the network one showing when it is live, and emb3r now turns down camera, location and notification requests by name rather than leaving them to a default it never chose.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card that means minutes for a greeting, and a stop button that looks broken because it is waiting on something that has barely begun. Gemini still only steps in for questions that look like they need current information; a provider you chose by hand and typed a key for now simply answers.
Replies did not say what answered them. The salamander tells you whether a reply came from this machine or from the web, which is the part that matters for privacy, but it cannot tell you which model. Replies now carry the name at the end — your local model, or the remote model and the host it came from.
The salamander beside each reply was invisible. It was drawing as a solid black shape on a nearly black background. Every rule that gives the creature its colours was written for the one in the header; the small copy was never included, so it fell back to plain black. It has its colours now.
The mark in the header was nearly invisible on the dark theme. Its outline was the same colour as the panel behind it, so a quarter of the drawing was rendering as background and the creature looked like it had pieces missing.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you w...
emb3r v1.36.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.36.0 — talking works in the installed app, and Macs can update again
Two of these are apologies.
Talking to Ember never worked in the installed app. It worked here while it was being built, and it worked in every test — but the version you download was missing one piece the speech recogniser loads before it does anything, so holding the button and speaking always ended in an error about a missing package. It has been broken since voice shipped, on Windows and on Mac. It is included now, and this time the check was done on a real installed copy rather than on the development one.
Updates install themselves on a Mac again. macOS refuses to let an application replace itself unless it was signed with an Apple certificate, which emb3r does not have — so the built-in updater was being turned away at the last step, after downloading the whole thing. emb3r now does that part itself: it downloads the new version, checks it against the checksum published alongside the release, and puts it in place. If anything fails at the final step the old version is put back rather than left half-replaced.
Worth knowing what that trade is. The Apple certificate is a promise that the app came from who it says; without one, the check is that the download matches the checksum the release published over an encrypted connection. That is a real check, and it is not the same check.
History, Talk and Settings are one menu now. They were three buttons crowding the corner with the network light — the one thing up there that is a promise rather than a control. The light has the room now.
v1.35.0 — a bigger face, and replies that are not waiting on the disk
Two things: Ember's face got much larger and learned five new expressions, and a memory bug that made replies crawl on a busy machine is fixed.
Press Face in the top bar and she fills the window. She has always had a set of expressions tied to what the app is actually doing — you have only ever seen them at the size of a line of text. Now they are the size of the window.
Five new ones, all about talking to her. Wide-eyed while the microphone is open, thoughtful while she works out what you said, puzzled when she could not make it out. If your microphone is missing or refused she tells you with her face rather than with an error. Before this, "listening" borrowed the surprised expression — the one view built for not reading the screen was showing you the wrong thing.
She takes your colour. The face is painted with a gradient built from whatever accent you have picked, lighter at the top and deeper at the bottom. If you have not picked one, she burns in emb3r's own fire instead of the interface green.
Replies could take minutes on a machine that was low on memory, and that is fixed. emb3r decides how much conversation to hold based on how much memory it can spare. It was working that out from how much memory your computer has in total, which assumes emb3r is the only thing running. On a laptop with everything else open it was asking for four and a half gigabytes when well under one was actually free, and Windows made up the difference by shuffling to disk. That is not slow thinking, it is a machine out of room — and it looks identical from the outside.
It now looks at what is genuinely free as well, and takes less when the machine is busy. If you have memory to spare nothing changes at all; the smaller window is only chosen when the alternative is swapping.
What went out. Face mode briefly had a drawn creature in it, with its own animation loop. Three designs were tried and none earned their place — one managed to be unsettling, another childish. That is a lot of machinery to maintain for something the faces were already doing well, so it is gone.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came out the same length to within eight milliseconds.
A word about waiting. Making speech is slower than saying it, so Ember banks a couple of seconds of audio before she starts. Ordinary replies then run without a break. A very long answer can still catch up with her and pause part-way through; that is arithmetic rather than a fault, and it is better than what it replaced, which was a five-second stall in the middle of a sentence.
The download is bigger — about 135 MB. That is the two speech models, and they are in there so that all of the above works the first time you open emb3r, with the offline lock switched on, having fetched nothing at all.
The microphone is only ever open while you hold the button, there is an indicator beside the network one showing when it is live, and emb3r now turns down camera, location and notification requests by name rather than leaving them to a default it never chose.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card that means minutes for a greeting, and a stop button that looks broken because it is waiting on something that has barely begun. Gemini still only steps in for questions that look like they need current information; a provider you chose by hand and typed a key for now simply answers.
Replies did not say what answered them. The salamander tells you whether a reply came from this machine or from the web, which is the part that matters for privacy, but it cannot tell you which model. Replies now carry the name at the end — your local model, or the remote model and the host it came from.
The salamander beside each reply was invisible. It was drawing as a solid black shape on a nearly black background. Every rule that gives the creature its colours was written for the one in the header; the small copy was never included, so it fell back to plain black. It has its colours now.
The mark in the header was nearly invisible on the dark theme. Its outline was the same colour as the panel behind it, so a quarter of the drawing was rendering as background and the creature looked like it had pieces missing.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you which ones it used. Any other provider is a model answering from what it already knows — often better or faster, but not the web. If you asked because you want today's news, keep Gemini.
Your key is treated the way the others are. It is written from the settings screen and never read back out, so the box stays empty afterwards even though the key is saved. The endpoint is checked before anything is sent to it: https only, or 127.0.0.1 if the server is on this machine, because anything else would put your key on the wire in the clear. A typo fails there and then instead of when you next ask a question.
The offline lock still applies to all of it, and the little indicator in the corner names the service you chose rather than saying "network activity".
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made...
emb3r v1.35.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.35.0 — a bigger face, and replies that are not waiting on the disk
Two things: Ember's face got much larger and learned five new expressions, and a memory bug that made replies crawl on a busy machine is fixed.
Press Face in the top bar and she fills the window. She has always had a set of expressions tied to what the app is actually doing — you have only ever seen them at the size of a line of text. Now they are the size of the window.
Five new ones, all about talking to her. Wide-eyed while the microphone is open, thoughtful while she works out what you said, puzzled when she could not make it out. If your microphone is missing or refused she tells you with her face rather than with an error. Before this, "listening" borrowed the surprised expression — the one view built for not reading the screen was showing you the wrong thing.
She takes your colour. The face is painted with a gradient built from whatever accent you have picked, lighter at the top and deeper at the bottom. If you have not picked one, she burns in emb3r's own fire instead of the interface green.
Replies could take minutes on a machine that was low on memory, and that is fixed. emb3r decides how much conversation to hold based on how much memory it can spare. It was working that out from how much memory your computer has in total, which assumes emb3r is the only thing running. On a laptop with everything else open it was asking for four and a half gigabytes when well under one was actually free, and Windows made up the difference by shuffling to disk. That is not slow thinking, it is a machine out of room — and it looks identical from the outside.
It now looks at what is genuinely free as well, and takes less when the machine is busy. If you have memory to spare nothing changes at all; the smaller window is only chosen when the alternative is swapping.
What went out. Face mode briefly had a drawn creature in it, with its own animation loop. Three designs were tried and none earned their place — one managed to be unsettling, another childish. That is a lot of machinery to maintain for something the faces were already doing well, so it is gone.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came out the same length to within eight milliseconds.
A word about waiting. Making speech is slower than saying it, so Ember banks a couple of seconds of audio before she starts. Ordinary replies then run without a break. A very long answer can still catch up with her and pause part-way through; that is arithmetic rather than a fault, and it is better than what it replaced, which was a five-second stall in the middle of a sentence.
The download is bigger — about 135 MB. That is the two speech models, and they are in there so that all of the above works the first time you open emb3r, with the offline lock switched on, having fetched nothing at all.
The microphone is only ever open while you hold the button, there is an indicator beside the network one showing when it is live, and emb3r now turns down camera, location and notification requests by name rather than leaving them to a default it never chose.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card that means minutes for a greeting, and a stop button that looks broken because it is waiting on something that has barely begun. Gemini still only steps in for questions that look like they need current information; a provider you chose by hand and typed a key for now simply answers.
Replies did not say what answered them. The salamander tells you whether a reply came from this machine or from the web, which is the part that matters for privacy, but it cannot tell you which model. Replies now carry the name at the end — your local model, or the remote model and the host it came from.
The salamander beside each reply was invisible. It was drawing as a solid black shape on a nearly black background. Every rule that gives the creature its colours was written for the one in the header; the small copy was never included, so it fell back to plain black. It has its colours now.
The mark in the header was nearly invisible on the dark theme. Its outline was the same colour as the panel behind it, so a quarter of the drawing was rendering as background and the creature looked like it had pieces missing.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you which ones it used. Any other provider is a model answering from what it already knows — often better or faster, but not the web. If you asked because you want today's news, keep Gemini.
Your key is treated the way the others are. It is written from the settings screen and never read back out, so the box stays empty afterwards even though the key is saved. The endpoint is checked before anything is sent to it: https only, or 127.0.0.1 if the server is on this machine, because anything else would put your key on the wire in the clear. A typo fails there and then instead of when you next ask a question.
The offline lock still applies to all of it, and the little indicator in the corner names the service you chose rather than saying "network activity".
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made Qwen3.5 unusable last week.
There is no longer a fixed number. When a model loads, emb3r asks that model what a larger window would cost, compares it against what the machine has spare, and takes the largest one that fits. On this laptop that is four times what it was.
Asking each model matters more than it sounds. A 16,384-token window costs 2,843MB with Llama 3.2 3B and 1,418MB with Qwen2.5 3B — the same setting, nearly twice the memory. One number for everything would have been too much for some machines and too cautious for the rest.
It can only go up. The old size is the floor, so a machine that cannot afford more gets exactly what it had before, and one that can gets more without being asked.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not ...
emb3r v1.34.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.34.0 — she can talk, and hear you
emb3r can be spoken to, and answers out loud. Both halves run on your own machine, which is the whole difficulty and most of the work.
Turn on Speech in Settings > Display and Ember reads her replies aloud. There is no voice picker, on purpose — she has one voice the way she has one face.
Press Face in the top bar to talk to her. Ember fills the window, the transcript and the typing box go away, and you hold the button or the spacebar while you speak. What she heard is printed on screen before she answers it, because being misheard is the one thing that goes wrong with talking to a computer and you should never have to guess whether that is what happened. Esc puts everything back.
None of it leaves this machine, and that took the long route. Every convenient way to add speech to an app like this sends audio to somebody's server — the recognition built into browsers streams your microphone to Google. Using it would have made the sentence on the front of this app false. So emb3r carries its own two speech models, one for speaking and one for listening, and both run here. The network indicator in the corner stays dark the whole time you are talking to her, which is the point.
The first attempt was thrown away. Windows will read text aloud using the voices it already has, and that version worked within an hour. It sounded like a machine reading a receipt, and there was no fixing it — this computer has eight voices installed and every one of them is the old kind. Its speed control turned out to be a fiction as well: set to 0.8, 1.0 or 1.2, the same sentence came out the same length to within eight milliseconds.
A word about waiting. Making speech is slower than saying it, so Ember banks a couple of seconds of audio before she starts. Ordinary replies then run without a break. A very long answer can still catch up with her and pause part-way through; that is arithmetic rather than a fault, and it is better than what it replaced, which was a five-second stall in the middle of a sentence.
The download is bigger — about 135 MB. That is the two speech models, and they are in there so that all of the above works the first time you open emb3r, with the offline lock switched on, having fetched nothing at all.
The microphone is only ever open while you hold the button, there is an indicator beside the network one showing when it is live, and emb3r now turns down camera, location and notification requests by name rather than leaving them to a default it never chose.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card that means minutes for a greeting, and a stop button that looks broken because it is waiting on something that has barely begun. Gemini still only steps in for questions that look like they need current information; a provider you chose by hand and typed a key for now simply answers.
Replies did not say what answered them. The salamander tells you whether a reply came from this machine or from the web, which is the part that matters for privacy, but it cannot tell you which model. Replies now carry the name at the end — your local model, or the remote model and the host it came from.
The salamander beside each reply was invisible. It was drawing as a solid black shape on a nearly black background. Every rule that gives the creature its colours was written for the one in the header; the small copy was never included, so it fell back to plain black. It has its colours now.
The mark in the header was nearly invisible on the dark theme. Its outline was the same colour as the panel behind it, so a quarter of the drawing was rendering as background and the creature looked like it had pieces missing.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you which ones it used. Any other provider is a model answering from what it already knows — often better or faster, but not the web. If you asked because you want today's news, keep Gemini.
Your key is treated the way the others are. It is written from the settings screen and never read back out, so the box stays empty afterwards even though the key is saved. The endpoint is checked before anything is sent to it: https only, or 127.0.0.1 if the server is on this machine, because anything else would put your key on the wire in the clear. A typo fails there and then instead of when you next ask a question.
The offline lock still applies to all of it, and the little indicator in the corner names the service you chose rather than saying "network activity".
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made Qwen3.5 unusable last week.
There is no longer a fixed number. When a model loads, emb3r asks that model what a larger window would cost, compares it against what the machine has spare, and takes the largest one that fits. On this laptop that is four times what it was.
Asking each model matters more than it sounds. A 16,384-token window costs 2,843MB with Llama 3.2 3B and 1,418MB with Qwen2.5 3B — the same setting, nearly twice the memory. One number for everything would have been too much for some machines and too cautious for the rest.
It can only go up. The old size is the floor, so a machine that cannot afford more gets exactly what it had before, and one that can gets more without being asked.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not finish in nine minutes on a laptop without a graphics card.
The honest part is how it got on the list. Three things were checked before it shipped: that the file existed, that it was the size it claimed, and that emb3r's engine recognises the kind of model it is. All three were true. Not one of them involved asking it a question, and nobody did. Checking that a model loads is not checking that it answers, and everything on that list is supposed to work on the machine reading it.
If you had it selected, emb3r moves you to a model that works the next time it starts, and says so. The 2.55GB you downloaded is left where it is — deleting somebody's file to tidy up our own mistake is not ours to do. Settings > Models will remove it if you want the space.
v1.31.0 — who is speaking
Both speakers in a conversation were the same colour to within a rounding error. Measured: your text and Ember's each stood out against the background at 17.4 and 15.3 to one, and against each other at 1.14 to one. There was also nothing between one message and the next, so a few exchanges ran together into a single block of text. The only thing telling you who was talking was the word at the start of the line, and you had to read it to find out.
Bubbles would have been the obvious fix and the wrong one — this is a terminal, and the prompt is how a terminal has always said who is speaking. So the prompt stays and the speaker changes. Ember's lines are led by the salamander itself, yours are still words. One of you is an animal and one is text, which is not a distinction you have to squint at. A reply that came from the web burns brighter, which is the same thing the fire on the mark already means when a connection is open.
The salamander in the transcript is the same drawing as the one in the header, not a copy, so the two can never drift apart. It holds still down there — twenty breathing creatures in a conversation would be a fidget rather than a signal.
Messages now have space between them and a coloured edge in the speaker's colour, and long lines stop at a comfortable width instead of running the full width of a wide window.
Two things that should have been ther...
emb3r v1.33.1
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.33.1 — the provider you picked, actually used
Four things, all reported by someone using Groq, and the first one explains most of the rest.
A provider other than Gemini was never used. The check that decides whether a question goes out to the web was still asking "is there a Gemini key?" — so if you had set up Groq and no Gemini key, the answer was no, every time, for every message. Everything went to the local model. On a laptop without a graphics card that means minutes for a greeting, and a stop button that looks broken because it is waiting on something that has barely begun. Gemini still only steps in for questions that look like they need current information; a provider you chose by hand and typed a key for now simply answers.
Replies did not say what answered them. The salamander tells you whether a reply came from this machine or from the web, which is the part that matters for privacy, but it cannot tell you which model. Replies now carry the name at the end — your local model, or the remote model and the host it came from.
The salamander beside each reply was invisible. It was drawing as a solid black shape on a nearly black background. Every rule that gives the creature its colours was written for the one in the header; the small copy was never included, so it fell back to plain black. It has its colours now.
The mark in the header was nearly invisible on the dark theme. Its outline was the same colour as the panel behind it, so a quarter of the drawing was rendering as background and the creature looked like it had pieces missing.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you which ones it used. Any other provider is a model answering from what it already knows — often better or faster, but not the web. If you asked because you want today's news, keep Gemini.
Your key is treated the way the others are. It is written from the settings screen and never read back out, so the box stays empty afterwards even though the key is saved. The endpoint is checked before anything is sent to it: https only, or 127.0.0.1 if the server is on this machine, because anything else would put your key on the wire in the clear. A typo fails there and then instead of when you next ask a question.
The offline lock still applies to all of it, and the little indicator in the corner names the service you chose rather than saying "network activity".
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made Qwen3.5 unusable last week.
There is no longer a fixed number. When a model loads, emb3r asks that model what a larger window would cost, compares it against what the machine has spare, and takes the largest one that fits. On this laptop that is four times what it was.
Asking each model matters more than it sounds. A 16,384-token window costs 2,843MB with Llama 3.2 3B and 1,418MB with Qwen2.5 3B — the same setting, nearly twice the memory. One number for everything would have been too much for some machines and too cautious for the rest.
It can only go up. The old size is the floor, so a machine that cannot afford more gets exactly what it had before, and one that can gets more without being asked.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not finish in nine minutes on a laptop without a graphics card.
The honest part is how it got on the list. Three things were checked before it shipped: that the file existed, that it was the size it claimed, and that emb3r's engine recognises the kind of model it is. All three were true. Not one of them involved asking it a question, and nobody did. Checking that a model loads is not checking that it answers, and everything on that list is supposed to work on the machine reading it.
If you had it selected, emb3r moves you to a model that works the next time it starts, and says so. The 2.55GB you downloaded is left where it is — deleting somebody's file to tidy up our own mistake is not ours to do. Settings > Models will remove it if you want the space.
v1.31.0 — who is speaking
Both speakers in a conversation were the same colour to within a rounding error. Measured: your text and Ember's each stood out against the background at 17.4 and 15.3 to one, and against each other at 1.14 to one. There was also nothing between one message and the next, so a few exchanges ran together into a single block of text. The only thing telling you who was talking was the word at the start of the line, and you had to read it to find out.
Bubbles would have been the obvious fix and the wrong one — this is a terminal, and the prompt is how a terminal has always said who is speaking. So the prompt stays and the speaker changes. Ember's lines are led by the salamander itself, yours are still words. One of you is an animal and one is text, which is not a distinction you have to squint at. A reply that came from the web burns brighter, which is the same thing the fire on the mark already means when a connection is open.
The salamander in the transcript is the same drawing as the one in the header, not a copy, so the two can never drift apart. It holds still down there — twenty breathing creatures in a conversation would be a fidget rather than a signal.
Messages now have space between them and a coloured edge in the speaker's colour, and long lines stop at a comfortable width instead of running the full width of a wide window.
Two things that should have been there from the start. Keyboard focus is now visible — there was nothing anywhere in the app showing which control you were on, which made it very hard to use without a mouse. And the accent colour has a Reset to default, matching the two that already existed for the personality and the model name; typing a very dark colour left you with a grey app and no obvious way back.
v1.30.1 — a light loading screen for a light theme
If you use light mode, every launch began with a full-screen black rectangle that turned white the moment the app finished loading. The loading screen had one background colour and it was black, and everything drawn on it — the wordmark, the torches, the little spirit that runs along them — was coloured for sitting on black.
It follows the theme now. The colours are not simply the dark ones lightened: an unlit torch has to stay faint against whatever is behind it, so it stays faint, while the wordmark goes the other way and turns dark, because letters waiting to catch fire should look like cold metal rather than something already burnt out.
The flash was a second problem hiding behind the first. The theme was being applied by the same file that draws the rest of the window, and that file runs last, so the loading screen had already appeared before anything knew which theme to use. It is now decided before the first frame is drawn.
v1.30.0 — the fire tells you what it is doing
The header had quietly become the main event. On a normal laptop window it took well over half the screen while the conversation got about an eighth of it. The face was large and on its own line, with mood and status stacked underneath.
The face and the mood share one line now, and the status text is not drawn at all. The chat has roughly three times the room it had.
The status did not disappear so much as change medium. The fire on the salamander's back has a pace for each thing the machine is doing: slow when idle, quicker while it is reading a file, quicker still while it is thinking, and a hard flare when something is genuinely going out to the web. The offline lock banks it down to embers.
That last one is the reason for the whole change. A flare at the edge of your vision is a better way to know a connection opened than a small label you stopped reading weeks ago. Nothing new was drawn for it; each state only changes the tempo of the breathing that arrived in v1.28.0. The written status is still there for screen readers, because a burning coil says nothing to one.
Accent colours can now be typed as a hex code...
emb3r v1.33.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.33.0 — bring your own provider
Web access meant Gemini or nothing. It now also speaks the format almost every other service uses, so Groq, OpenRouter, Together, DeepSeek, Mistral and a server running on your own network all work. Settings > Web access asks for three things: the endpoint, your key, and the model name exactly as they write it.
One difference is worth being clear about, and it is written into the panel rather than left to be discovered. Gemini is set up to read live web pages and tell you which ones it used. Any other provider is a model answering from what it already knows — often better or faster, but not the web. If you asked because you want today's news, keep Gemini.
Your key is treated the way the others are. It is written from the settings screen and never read back out, so the box stays empty afterwards even though the key is saved. The endpoint is checked before anything is sent to it: https only, or 127.0.0.1 if the server is on this machine, because anything else would put your key on the wire in the clear. A typo fails there and then instead of when you next ask a question.
The offline lock still applies to all of it, and the little indicator in the corner names the service you chose rather than saying "network activity".
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made Qwen3.5 unusable last week.
There is no longer a fixed number. When a model loads, emb3r asks that model what a larger window would cost, compares it against what the machine has spare, and takes the largest one that fits. On this laptop that is four times what it was.
Asking each model matters more than it sounds. A 16,384-token window costs 2,843MB with Llama 3.2 3B and 1,418MB with Qwen2.5 3B — the same setting, nearly twice the memory. One number for everything would have been too much for some machines and too cautious for the rest.
It can only go up. The old size is the floor, so a machine that cannot afford more gets exactly what it had before, and one that can gets more without being asked.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not finish in nine minutes on a laptop without a graphics card.
The honest part is how it got on the list. Three things were checked before it shipped: that the file existed, that it was the size it claimed, and that emb3r's engine recognises the kind of model it is. All three were true. Not one of them involved asking it a question, and nobody did. Checking that a model loads is not checking that it answers, and everything on that list is supposed to work on the machine reading it.
If you had it selected, emb3r moves you to a model that works the next time it starts, and says so. The 2.55GB you downloaded is left where it is — deleting somebody's file to tidy up our own mistake is not ours to do. Settings > Models will remove it if you want the space.
v1.31.0 — who is speaking
Both speakers in a conversation were the same colour to within a rounding error. Measured: your text and Ember's each stood out against the background at 17.4 and 15.3 to one, and against each other at 1.14 to one. There was also nothing between one message and the next, so a few exchanges ran together into a single block of text. The only thing telling you who was talking was the word at the start of the line, and you had to read it to find out.
Bubbles would have been the obvious fix and the wrong one — this is a terminal, and the prompt is how a terminal has always said who is speaking. So the prompt stays and the speaker changes. Ember's lines are led by the salamander itself, yours are still words. One of you is an animal and one is text, which is not a distinction you have to squint at. A reply that came from the web burns brighter, which is the same thing the fire on the mark already means when a connection is open.
The salamander in the transcript is the same drawing as the one in the header, not a copy, so the two can never drift apart. It holds still down there — twenty breathing creatures in a conversation would be a fidget rather than a signal.
Messages now have space between them and a coloured edge in the speaker's colour, and long lines stop at a comfortable width instead of running the full width of a wide window.
Two things that should have been there from the start. Keyboard focus is now visible — there was nothing anywhere in the app showing which control you were on, which made it very hard to use without a mouse. And the accent colour has a Reset to default, matching the two that already existed for the personality and the model name; typing a very dark colour left you with a grey app and no obvious way back.
v1.30.1 — a light loading screen for a light theme
If you use light mode, every launch began with a full-screen black rectangle that turned white the moment the app finished loading. The loading screen had one background colour and it was black, and everything drawn on it — the wordmark, the torches, the little spirit that runs along them — was coloured for sitting on black.
It follows the theme now. The colours are not simply the dark ones lightened: an unlit torch has to stay faint against whatever is behind it, so it stays faint, while the wordmark goes the other way and turns dark, because letters waiting to catch fire should look like cold metal rather than something already burnt out.
The flash was a second problem hiding behind the first. The theme was being applied by the same file that draws the rest of the window, and that file runs last, so the loading screen had already appeared before anything knew which theme to use. It is now decided before the first frame is drawn.
v1.30.0 — the fire tells you what it is doing
The header had quietly become the main event. On a normal laptop window it took well over half the screen while the conversation got about an eighth of it. The face was large and on its own line, with mood and status stacked underneath.
The face and the mood share one line now, and the status text is not drawn at all. The chat has roughly three times the room it had.
The status did not disappear so much as change medium. The fire on the salamander's back has a pace for each thing the machine is doing: slow when idle, quicker while it is reading a file, quicker still while it is thinking, and a hard flare when something is genuinely going out to the web. The offline lock banks it down to embers.
That last one is the reason for the whole change. A flare at the edge of your vision is a better way to know a connection opened than a small label you stopped reading weeks ago. Nothing new was drawn for it; each state only changes the tempo of the breathing that arrived in v1.28.0. The written status is still there for screen readers, because a burning coil says nothing to one.
Accent colours can now be typed as a hex code, next to the wheel, since a wheel cannot land on an exact value. Anything that is not six hex digits is refused. Anything that would be unreadable against your theme is adjusted, and it says so rather than quietly handing you something else — a near-black accent on the dark theme is lifted until it can be read, and it tells you it lifted it.
Ember also knows who made it. Ask who built it and it will say Ziyan Dobaria, rather than naming whoever trained the model underneath or inventing a company.
v1.29.1 — the switch, beside its label
The checkbox under Settings > Memory was at one end of the panel and the words explaining it were at the other, with a hand's width of nothing in between. Text boxes in Settings are meant to run the full width; a checkbox is not, and this one had been left out of the rule that says so.
It was found by photographing the tab for the project's evidence document, not by reading the code. Nothing about the markup looks wrong, which is how it got through a review and a release.
v1.29.0 — it remembers, if you let it
Ember can be told things worth keeping between conversations. What you are working on, how you like your answers, the name of your dog. Settings > Memory. They live on this machine in the same file as the rest of your settings, they belong to whichever profile you are using, and a switch turns the whole thing off without deleting any of them.
The interesting part is what a memory is allowed to cost. The first version put every remembered fact into every single reply, which sounds harmless until you count it: twenty of the...
emb3r v1.32.0
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.32.0 — as much conversation as your machine can hold
emb3r could only hold about 4,096 tokens of conversation at a time — roughly three thousand words, counting everything: the instructions it runs on, anything you attached, and the reply it is in the middle of writing. Past that it starts dropping the oldest part, and if it cannot drop enough it stops and says so.
That number was doing real damage. Six extracts from a PDF came to 1,907 tokens on their own, leaving 187 for the answer, and the answer was cut off mid-sentence. It is also what made Qwen3.5 unusable last week.
There is no longer a fixed number. When a model loads, emb3r asks that model what a larger window would cost, compares it against what the machine has spare, and takes the largest one that fits. On this laptop that is four times what it was.
Asking each model matters more than it sounds. A 16,384-token window costs 2,843MB with Llama 3.2 3B and 1,418MB with Qwen2.5 3B — the same setting, nearly twice the memory. One number for everything would have been too much for some machines and too cautious for the rest.
It can only go up. The old size is the floor, so a machine that cannot afford more gets exactly what it had before, and one that can gets more without being asked.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not finish in nine minutes on a laptop without a graphics card.
The honest part is how it got on the list. Three things were checked before it shipped: that the file existed, that it was the size it claimed, and that emb3r's engine recognises the kind of model it is. All three were true. Not one of them involved asking it a question, and nobody did. Checking that a model loads is not checking that it answers, and everything on that list is supposed to work on the machine reading it.
If you had it selected, emb3r moves you to a model that works the next time it starts, and says so. The 2.55GB you downloaded is left where it is — deleting somebody's file to tidy up our own mistake is not ours to do. Settings > Models will remove it if you want the space.
v1.31.0 — who is speaking
Both speakers in a conversation were the same colour to within a rounding error. Measured: your text and Ember's each stood out against the background at 17.4 and 15.3 to one, and against each other at 1.14 to one. There was also nothing between one message and the next, so a few exchanges ran together into a single block of text. The only thing telling you who was talking was the word at the start of the line, and you had to read it to find out.
Bubbles would have been the obvious fix and the wrong one — this is a terminal, and the prompt is how a terminal has always said who is speaking. So the prompt stays and the speaker changes. Ember's lines are led by the salamander itself, yours are still words. One of you is an animal and one is text, which is not a distinction you have to squint at. A reply that came from the web burns brighter, which is the same thing the fire on the mark already means when a connection is open.
The salamander in the transcript is the same drawing as the one in the header, not a copy, so the two can never drift apart. It holds still down there — twenty breathing creatures in a conversation would be a fidget rather than a signal.
Messages now have space between them and a coloured edge in the speaker's colour, and long lines stop at a comfortable width instead of running the full width of a wide window.
Two things that should have been there from the start. Keyboard focus is now visible — there was nothing anywhere in the app showing which control you were on, which made it very hard to use without a mouse. And the accent colour has a Reset to default, matching the two that already existed for the personality and the model name; typing a very dark colour left you with a grey app and no obvious way back.
v1.30.1 — a light loading screen for a light theme
If you use light mode, every launch began with a full-screen black rectangle that turned white the moment the app finished loading. The loading screen had one background colour and it was black, and everything drawn on it — the wordmark, the torches, the little spirit that runs along them — was coloured for sitting on black.
It follows the theme now. The colours are not simply the dark ones lightened: an unlit torch has to stay faint against whatever is behind it, so it stays faint, while the wordmark goes the other way and turns dark, because letters waiting to catch fire should look like cold metal rather than something already burnt out.
The flash was a second problem hiding behind the first. The theme was being applied by the same file that draws the rest of the window, and that file runs last, so the loading screen had already appeared before anything knew which theme to use. It is now decided before the first frame is drawn.
v1.30.0 — the fire tells you what it is doing
The header had quietly become the main event. On a normal laptop window it took well over half the screen while the conversation got about an eighth of it. The face was large and on its own line, with mood and status stacked underneath.
The face and the mood share one line now, and the status text is not drawn at all. The chat has roughly three times the room it had.
The status did not disappear so much as change medium. The fire on the salamander's back has a pace for each thing the machine is doing: slow when idle, quicker while it is reading a file, quicker still while it is thinking, and a hard flare when something is genuinely going out to the web. The offline lock banks it down to embers.
That last one is the reason for the whole change. A flare at the edge of your vision is a better way to know a connection opened than a small label you stopped reading weeks ago. Nothing new was drawn for it; each state only changes the tempo of the breathing that arrived in v1.28.0. The written status is still there for screen readers, because a burning coil says nothing to one.
Accent colours can now be typed as a hex code, next to the wheel, since a wheel cannot land on an exact value. Anything that is not six hex digits is refused. Anything that would be unreadable against your theme is adjusted, and it says so rather than quietly handing you something else — a near-black accent on the dark theme is lifted until it can be read, and it tells you it lifted it.
Ember also knows who made it. Ask who built it and it will say Ziyan Dobaria, rather than naming whoever trained the model underneath or inventing a company.
v1.29.1 — the switch, beside its label
The checkbox under Settings > Memory was at one end of the panel and the words explaining it were at the other, with a hand's width of nothing in between. Text boxes in Settings are meant to run the full width; a checkbox is not, and this one had been left out of the rule that says so.
It was found by photographing the tab for the project's evidence document, not by reading the code. Nothing about the markup looks wrong, which is how it got through a review and a release.
v1.29.0 — it remembers, if you let it
Ember can be told things worth keeping between conversations. What you are working on, how you like your answers, the name of your dog. Settings > Memory. They live on this machine in the same file as the rest of your settings, they belong to whichever profile you are using, and a switch turns the whole thing off without deleting any of them.
The interesting part is what a memory is allowed to cost. The first version put every remembered fact into every single reply, which sounds harmless until you count it: twenty of them is about a quarter of everything the model can hold, spent before you have typed a word. On a laptop with no room to spare that is the difference between an answer and a truncated one.
So they work the way attached files already do here. Only the ones your question actually touches get sent, and they are dropped again afterwards. Ask about your dog and it brings up the dog. Ask about anything else and the dog costs nothing. In practice that is around one percent of the window instead of a quarter.
Getting the matching right needed one correction. Asking what is my dog called also pulled up "an Electron app called emb3r" — the two share the word "called". A list of words to ignore would never have caught that, because "called" is an ordinary word right up until you have written it twice. Ember now works out for itself which words are too common across your own memories to tell them apart.
Also new: Qwen3.5 4B, the newest model on the list and the quickest of the capable ones, at 2.6GB. Checked before it was offered — that it exists, that it is the size claimed, and that the engine inside emb3r actually knows how to load it. A model that fails when you first ask it something is worse than one that was never listed.
v1.28.0 — the salamander breath...
emb3r v1.31.1
A small terminal-dwelling AI companion that runs a language model entirely on your own machine. By default, nothing you type is sent to a server — and as of v1.1.0 you can verify that, and enforce it.
v1.31.1 — taking a model back off the shelf
Qwen3.5 4B was added to the model list in v1.29.0 and is being removed. If you tried it, you will have seen every reply end in a message about failing to compress the chat history. That was not your machine.
It is a reasoning model: it works through an answer to itself before writing one, and that working-out is hidden from you but still has to be held in memory. There is only room for about 4,096 tokens of conversation, and the thinking filled it before anything was said. Once full, emb3r tries to make room by dropping the oldest part of the conversation, and it cannot drop the instructions it needs to keep — so it gives up and says so. Asked to say "hi", the model produced forty tokens of private reasoning and not one visible character.
It was not fast either, which is the other reason it was there. A single long reply did not finish in nine minutes on a laptop without a graphics card.
The honest part is how it got on the list. Three things were checked before it shipped: that the file existed, that it was the size it claimed, and that emb3r's engine recognises the kind of model it is. All three were true. Not one of them involved asking it a question, and nobody did. Checking that a model loads is not checking that it answers, and everything on that list is supposed to work on the machine reading it.
If you had it selected, emb3r moves you to a model that works the next time it starts, and says so. The 2.55GB you downloaded is left where it is — deleting somebody's file to tidy up our own mistake is not ours to do. Settings > Models will remove it if you want the space.
v1.31.0 — who is speaking
Both speakers in a conversation were the same colour to within a rounding error. Measured: your text and Ember's each stood out against the background at 17.4 and 15.3 to one, and against each other at 1.14 to one. There was also nothing between one message and the next, so a few exchanges ran together into a single block of text. The only thing telling you who was talking was the word at the start of the line, and you had to read it to find out.
Bubbles would have been the obvious fix and the wrong one — this is a terminal, and the prompt is how a terminal has always said who is speaking. So the prompt stays and the speaker changes. Ember's lines are led by the salamander itself, yours are still words. One of you is an animal and one is text, which is not a distinction you have to squint at. A reply that came from the web burns brighter, which is the same thing the fire on the mark already means when a connection is open.
The salamander in the transcript is the same drawing as the one in the header, not a copy, so the two can never drift apart. It holds still down there — twenty breathing creatures in a conversation would be a fidget rather than a signal.
Messages now have space between them and a coloured edge in the speaker's colour, and long lines stop at a comfortable width instead of running the full width of a wide window.
Two things that should have been there from the start. Keyboard focus is now visible — there was nothing anywhere in the app showing which control you were on, which made it very hard to use without a mouse. And the accent colour has a Reset to default, matching the two that already existed for the personality and the model name; typing a very dark colour left you with a grey app and no obvious way back.
v1.30.1 — a light loading screen for a light theme
If you use light mode, every launch began with a full-screen black rectangle that turned white the moment the app finished loading. The loading screen had one background colour and it was black, and everything drawn on it — the wordmark, the torches, the little spirit that runs along them — was coloured for sitting on black.
It follows the theme now. The colours are not simply the dark ones lightened: an unlit torch has to stay faint against whatever is behind it, so it stays faint, while the wordmark goes the other way and turns dark, because letters waiting to catch fire should look like cold metal rather than something already burnt out.
The flash was a second problem hiding behind the first. The theme was being applied by the same file that draws the rest of the window, and that file runs last, so the loading screen had already appeared before anything knew which theme to use. It is now decided before the first frame is drawn.
v1.30.0 — the fire tells you what it is doing
The header had quietly become the main event. On a normal laptop window it took well over half the screen while the conversation got about an eighth of it. The face was large and on its own line, with mood and status stacked underneath.
The face and the mood share one line now, and the status text is not drawn at all. The chat has roughly three times the room it had.
The status did not disappear so much as change medium. The fire on the salamander's back has a pace for each thing the machine is doing: slow when idle, quicker while it is reading a file, quicker still while it is thinking, and a hard flare when something is genuinely going out to the web. The offline lock banks it down to embers.
That last one is the reason for the whole change. A flare at the edge of your vision is a better way to know a connection opened than a small label you stopped reading weeks ago. Nothing new was drawn for it; each state only changes the tempo of the breathing that arrived in v1.28.0. The written status is still there for screen readers, because a burning coil says nothing to one.
Accent colours can now be typed as a hex code, next to the wheel, since a wheel cannot land on an exact value. Anything that is not six hex digits is refused. Anything that would be unreadable against your theme is adjusted, and it says so rather than quietly handing you something else — a near-black accent on the dark theme is lifted until it can be read, and it tells you it lifted it.
Ember also knows who made it. Ask who built it and it will say Ziyan Dobaria, rather than naming whoever trained the model underneath or inventing a company.
v1.29.1 — the switch, beside its label
The checkbox under Settings > Memory was at one end of the panel and the words explaining it were at the other, with a hand's width of nothing in between. Text boxes in Settings are meant to run the full width; a checkbox is not, and this one had been left out of the rule that says so.
It was found by photographing the tab for the project's evidence document, not by reading the code. Nothing about the markup looks wrong, which is how it got through a review and a release.
v1.29.0 — it remembers, if you let it
Ember can be told things worth keeping between conversations. What you are working on, how you like your answers, the name of your dog. Settings > Memory. They live on this machine in the same file as the rest of your settings, they belong to whichever profile you are using, and a switch turns the whole thing off without deleting any of them.
The interesting part is what a memory is allowed to cost. The first version put every remembered fact into every single reply, which sounds harmless until you count it: twenty of them is about a quarter of everything the model can hold, spent before you have typed a word. On a laptop with no room to spare that is the difference between an answer and a truncated one.
So they work the way attached files already do here. Only the ones your question actually touches get sent, and they are dropped again afterwards. Ask about your dog and it brings up the dog. Ask about anything else and the dog costs nothing. In practice that is around one percent of the window instead of a quarter.
Getting the matching right needed one correction. Asking what is my dog called also pulled up "an Electron app called emb3r" — the two share the word "called". A list of words to ignore would never have caught that, because "called" is an ordinary word right up until you have written it twice. Ember now works out for itself which words are too common across your own memories to tell them apart.
Also new: Qwen3.5 4B, the newest model on the list and the quickest of the capable ones, at 2.6GB. Checked before it was offered — that it exists, that it is the size claimed, and that the engine inside emb3r actually knows how to load it. A model that fails when you first ask it something is worse than one that was never listed.
v1.28.0 — the salamander breathes
The mark in the header stood still. It breathes now: each of the four flames on the creature's back lifts from its own base and settles, and the hottest part of the fire flickers on a faster clock inside that slower breath. The four run on lengths that do not divide into each other, so they drift apart instead of pulsing in unison. The body does not move — the coil is a ring, and a ring that breathes reads as a wobble.
It moves in whole pixels rather than gliding. Smoothing the motion was tried and thrown away: it did glide, but it also made the creature look soft, and softness is the one thing a pixel-art animal cannot afford.
Underneath, the drawing is no longer 418 separate squares. Each shade is now a single shape, which is what made animating it possible at all — squares laid edge to edge crack apart along their shared edges the moment they stop being snapped to whole pixels. 418 pieces became 17.
If your system is set to reduce motion, none of this runs.
v1.27.3 — the coil comes inside
The coiled salamander was drawn for the taskbar and stayed there; the header kept the long side-on drawing. It is the mark in the application now too, square beside the six-line wordmark rather than running most of its width.
The colouring changed with...