Problem
Currently, even if you have the 'Faster first result' off (on AI Enhancement → Fluid Intelligence), once you've used the model, it then sits in memory until the app is closed.
This is obviously perfect if you use dictation super regularly as it keeps the model warm... but for me (and I assume other users), I might dictate a few hundred words, then nothing for the next hour or so. During that period of inactivity, the fluid-intelligence-mlx process is running in the background taking ~3GB memory.
Proposed solution
To have an additional toggle in the Fluid Intelligence Edit Provider settings that allows the app to drops the model from memory after 3-5 mins of inactivity.
I know that for me, I'd rather have the extra memory on hand (or lower resting memory pressure with all my Chrome tabs, lol) in exchange for a slower cold-start at the beginning of a dictation session. Maybe the inactivity window could be user-customised?
Alternatives considered
No response
Problem
Currently, even if you have the 'Faster first result' off (on AI Enhancement → Fluid Intelligence), once you've used the model, it then sits in memory until the app is closed.
This is obviously perfect if you use dictation super regularly as it keeps the model warm... but for me (and I assume other users), I might dictate a few hundred words, then nothing for the next hour or so. During that period of inactivity, the
fluid-intelligence-mlxprocess is running in the background taking ~3GB memory.Proposed solution
To have an additional toggle in the Fluid Intelligence Edit Provider settings that allows the app to drops the model from memory after 3-5 mins of inactivity.
I know that for me, I'd rather have the extra memory on hand (or lower resting memory pressure with all my Chrome tabs, lol) in exchange for a slower cold-start at the beginning of a dictation session. Maybe the inactivity window could be user-customised?
Alternatives considered
No response