I was looking back at some old lemmee posts and came across GPT4All. Didn’t get much sleep last night as it’s awesome, even on my old (10yo) laptop with a Compute 5.0 NVidia card.
Still, I’m after more, I’d like to be able to get image creation and view it in the conversation, if it generates python code, to be able to run it (I’m using Debian, and have a default python env set up). Local file analysis also useful. CUDA Compute 5.0 / vulkan compatibility needed too with the option to use some of the smaller models (1-3B for example). Also a local API would be nice for my own python experiments.
Is there anything that can tick the boxes? Even if I have to scoot across models for some of the features? I’d prefer more of a desktop client application than a docker container running in the background.
Ollama for API, which you can integrate into Open WebUI. You can also integrate image generation with ComfyUI I believe.
It’s less of a hassle to use Docker for Open WebUI, but ollama works as a regular CLI tool.
ChainLit is a super ez UI too. Ollama works well with Semantic Kernal (for integration with existing code) and langChain (for agent orchestration). I’m working on building MCP interaction with ComfyUI’s API, it’s a pain in the ass.
But won’t this be a mish-mash of different docker containers and projects creating an installation, dependency, upgrade nightmare?
All the ones I mentioned can be installed with pip or uv if I am not mistaken. It would probably be more finicky than containers that you can put behind a reverse proxy, but it is possible if you wish to go that route. Ollama will also run system-wide, so any project will be able to use its API without you having to create a separate environment and download the same model twice in order to use it.
This is what I do its excellent.
Maybe LocalAI? It doesn’t do python code execution, but pretty much all of the rest.
This looks interesting - do you have experience of it? How reliable / efficient is it?
LocalAI is pretty good but resource-intensive. I ran it on a vps in the past.
I’ve discovered jan.ai which is far faster than GPT4All, and visually a little nicer.
EDIT: After using it for an hour or so, it seems to crash all the time, I keep on having to reset it, and currently am facing it freezing for no reason.
I also started using this recently and it’s very plug and play. Just open and run. It’s the only client so far that feels like I could recommend to non-geeks.
AUTOMATIC1111?
The main limitation is the VRAM, but I doubt any model is going to be particularly fast.
I think
phi3:mini
on ollama might be an okish fit for python, since it’s a small model, but was trained on python codebases.I’m getting very-near real-time on my old laptop. Maybe a delay of 1-2s whilst it creates the response
You should try https://cherry-ai.com/ … It’s the most advanced client out there. I personally use Ollama for running the models and Mistral API for advnaced tasks.
But its website is Chinese. Also what’s the github?
It’s fully open source and free (as in beer).
Would like fries or a jetpack with that?
You should try https://cherry-ai.com/ … It’s the most advanced client out there. I personally use Ollama for running the models and Mistral API for advnaced tasks.