Replies: 9 comments 9 replies
|
Let's discuss this more to understand what everyone thinks about this and and if a flag for inactivity is good enough. Let's sit on this for a lil more and see where to take this! Thanks for raising this issue. |
|
First off, Fluid Voice is incredible — seriously, great work. It’s very quickly become a vital part of certain parts of my workflow. I think the key distinction here is “always ready” vs “ready when I need it.” For someone dictating constantly, keeping the model warm is absolutely the right behaviour. But for people like me who use FluidVoice in bursts, it means I can dictate for 2 minutes and then effectively pay a ~3GB RAM tax for the next hour while I’m back in Chrome, Cursor, Canva, etc. And that’s the slightly painful bit — FluidVoice is brilliant precisely because it disappears into my workflow. But seeing memory pressure climb afterwards makes me start thinking about FluidVoice again, and occasionally whether I should quit it to reclaim the RAM. I’d happily accept a 2–3 second cold start on my next dictation in exchange for getting those resources back during the 95% of the time I’m not speaking. I actually think an inactivity option could make everyone happy: Keep Fluid Intelligence loaded Always Or even just an initial “Unload after inactivity” flag with a sensible fixed timeout would be enough to test whether people use it. The reason I’d love to see this tried is that it removes one of the very few compromises I currently feel I’m making to use FluidVoice. It would turn Fluid Intelligence from something I sometimes need to manage, into something I never have to think about. For anyone else using local Fluid Intelligence: would you rather keep the instant response, or reclaim the ~3GB after you stop dictating? |
|
Would a 1GB model cause this too? :)
…On Fri, Aug 21, 2026 at 6:52 AM Dan T ***@***.***> wrote:
First off, Fluid Voice is incredible — seriously, great work. It’s very
quickly become a vital part of certain parts of my workflow.
I think the key *distinction here is “always ready” vs “ready when I need
it.”*
For someone dictating constantly, keeping the model warm is absolutely the
right behaviour. But for people like me who use FluidVoice in bursts, it
means I can dictate for 2 minutes and then effectively pay *a ~3GB RAM
tax for the next hour* while I’m back in Chrome, Cursor, Canva, etc.
And that’s the slightly painful bit — FluidVoice is brilliant precisely
because it disappears into my workflow. But seeing memory pressure climb
afterwards makes me start thinking about FluidVoice again, and occasionally
whether I should quit it to reclaim the RAM.
I’d happily accept a 2–3 second cold start on my next dictation in
exchange for getting those resources back during the 95% of the time I’m
not speaking.
I actually think an inactivity option could make everyone happy:
Keep Fluid Intelligence loaded
Always
5 minutes
15 minutes
30 minutes
Or even just an initial “Unload after inactivity” flag with a sensible
fixed timeout would be enough to test whether people use it.
The reason I’d love to see this tried is that it removes one of the very
few compromises I currently feel I’m making to use FluidVoice. It would
turn Fluid Intelligence from something I sometimes need to manage, into
something I never have to think about.
For anyone else using local Fluid Intelligence: would you rather keep the
instant response, or reclaim the ~3GB after you stop dictating?
—
Reply to this email directly, view it on GitHub
<#854?email_source=notifications&email_token=BVSOW2WSYDS2R63NAGBZT635LBHZZA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBRGA3DQMZRUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18106831>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BVSOW2ROFB44YK5PSHGSK3D5LBHZZAVCNFSNUABJKJSXA33TNF2G64TZHMYTANRRGMZDOMZRGE5UI2LTMN2XG43JN5XDWMJQGYYDQNJWGWQXMAQ>
.
You are receiving this because you commented.Message ID:
***@***.***>
|
|
Well, this one is currently better that the fluid-1 😉
…On Sat, Aug 22, 2026 at 8:44 AM Josh Dale ***@***.***> wrote:
Probably not, would need to try. But you wouldn’t be able to use the best
model!
—
Reply to this email directly, view it on GitHub
<#854?email_source=notifications&email_token=BVSOW2TIK754QGTT27FV6J35LG5U3A5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBRGE3TENRXUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18117267>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BVSOW2TXNZOTC6WSIKLMGRD5LG5U3AVCNFSNUABJKJSXA33TNF2G64TZHMYTANRRGMZDOMZRGE5UI2LTMN2XG43JN5XDWMJQGYYDQNJWGWQXMAQ>
.
You are receiving this because you commented.Message ID:
***@***.***>
|
|
Is there a way in which I can fist unload the enhancement model on FluidVoice (and free that 3+ GB RAM) and then try to access a locally hosted model (e.g. via llama, mlx-serve) by adding model server base url and other details of Settings in "Add Customer Provider"? |
|
Would you like to try out a 1.4gb model I have? Planning to release it soon
but would love for you to try and see if it still bothers you
…On Sun, Aug 30, 2026 at 3:32 PM EnderonNZ ***@***.***> wrote:
Would love to see this, my only barrier is the RAM the current fluid model
takes up.
—
Reply to this email directly, view it on GitHub
<#854?email_source=notifications&email_token=BVSOW2QVHLJIPQBH5LJ7OYT5MSTRHA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBSGEYDMNBTUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18210643>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BVSOW2U3VATCDLW3NRPEHE35MSTRHAVCNFSNUABJKJSXA33TNF2G64TZHMYTANRRGMZDOMZRGE5UI2LTMN2XG43JN5XDWMJQGYYDQNJWGWQXMAQ>
.
You are receiving this because you commented.Message ID:
***@***.***>
|
|
The enhancement model does not have to run on your computer (it can be on a
server for eg ) most of the time if you don't want to. If you still want
to, you can have different models for dictation, emails, edits, even, cmd
mode. So the options are plenty. You don't have to have one model for
everything, right?
Thats why you have an option for all
…On Mon, Aug 31, 2026 at 2:14 PM justauserid ***@***.***> wrote:
Would you like to try out a 1.4gb model I have? Planning to release it
soon but would love for you to try and see if it still bothers you
The reason I want to do is I want to try some models I have and the
hosting setups I am trying. It's just a 16GB RAM. With apps and everything
it doesn't leave a lot of RAM to begin with.
What I mean is, is this not odd that there's a loaded model for
enhancement and an option to connect with another model (via API) to do the
enhancement at the same time w/o unloading the enhancement model?
—
Reply to this email directly, view it on GitHub
<#854?email_source=notifications&email_token=BVSOW2X5KTZQEKYN6ZVLWIT5MXTEPA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBSGI2DEMJYUZZGKYLTN5XKOY3PNVWWK3TUUVSXMZLOOSWGM33PORSXEX3DNRUWG2Y#discussioncomment-18224218>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/BVSOW2R3KPMX6YNZ4XEW2JL5MXTEPAVCNFSNUABJKJSXA33TNF2G64TZHMYTANRRGMZDOMZRGE5UI2LTMN2XG43JN5XDWMJQGYYDQNJWGWQXMAQ>
.
You are receiving this because you commented.Message ID:
***@***.***>
|
|
@justauserid - Thanks for explaining - that did it land well last time. Let's try this - Click edit and remove it from and reset it like shown below. lmk if it still takes the RAM. I will fix it in the next update if so.
Tbh - all other models are independent of what I ship - you can technically have Ollama, MLX, LMS and all of that loaded and they won't talk to each other just like FI right here. It's overall a fair ask but that's why I have it as 'Verified' to load them as fast as possible for users. The delay would kill the UX and I do not want that, unfortunately. So, coming back to our discussion - will you please give that a go to see if it removes it from memory? if not, it's a bug on my side and I will fix it going forward. (Also, FWIW - I am going to ship that same model soon at 1/5x - 2x the speed you'd get rn with mlx-serve very soon :) ) Thanks for understanding, and let me know how it goes. I'm looking at this issue carefully :) |
|
Totally agree with everything you say. I use local models and it works well ( LMStudio) - you need to select it on verified (Maybe bad UX? idk,, will have to think through it as its very complicated to give it all but still make it easy to understand haha) - 4GB was not my goal to ship a custom model as I know how important RAM is - my model will probably be better than others for most cases and I am so close to wrapping up a lot of smaller models ( 0.2B to 2B) models so the model size will be < 1.5GB for almost all of them and incredibly faster. that's the goal! I will add the toggle soon tho. |



Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
Currently, even if you have the 'Faster first result' off (on AI Enhancement → Fluid Intelligence), once you've used the model, it then sits in memory until the app is closed.
This is obviously perfect if you use dictation super regularly as it keeps the model warm... but for me (and I assume other users), I might dictate a few hundred words, then nothing for the next hour or so. During that period of inactivity, the
fluid-intelligence-mlxprocess is running in the background taking ~3GB memory.Proposed solution
To have an additional toggle in the Fluid Intelligence Edit Provider settings that allows the app to drops the model from memory after 3-5 mins of inactivity.
I know that for me, I'd rather have the extra memory on hand (or lower resting memory pressure with all my Chrome tabs, lol) in exchange for a slower cold-start at the beginning of a dictation session. Maybe the inactivity window could be user-customised?
Alternatives considered
No response
All reactions