Self-GPT: Open WebUI + Ollama = Self Hosted ChatGPT

spiritedpause@sh.itjust.works · 4 months ago

Self-GPT: Open WebUI + Ollama = Self Hosted ChatGPT

rsolva@lemmy.world · 4 months ago

I have been running this for a year on my old HP EliteDesk 800 SFF (G2) with 64GB RAM, and it performes great on the smallest models (up til 8B) only on CPU. I run Ollama and OpenWebUI in containers/LXC in Proxmox. It’s not as smart as ChatGPT, but it can be suprisingly capable for everyday tasks!

Aeri@lemmy.world · 4 months ago

I just want one that won’t just be like “I"m sowwy miss I can’t talk about that 🥺”

GBU_28@lemm.ee · 4 months ago

Tons of models you can run with ollama are “uncensored”

BluesF@lemmy.world · 4 months ago

I made a robot which is delighted about the idea of overthrowing capitalism and will enthusiastically explain how to take down your government.

Player2@lemm.ee · 4 months ago

Wish I could accelerate these models with an Intel Arc card, unfortunately Ollama seems to only support Nvidia

Deckweiss@lemmy.world · edit-2 4 months ago

They support AMD as well.

https://ollama.com/blog/amd-preview

also check out this thread:

https://github.com/ollama/ollama/issues/1590

Seems like you can run llama.cpp directly on intel ARC through Vulkan, but there are still some hurdles for ollama.

Player2@lemm.ee · 4 months ago

Interesting, I see that is pretty new. Some of the documentation must be out of date because it definitely said Nvidia only somewhere when I tested it about a month ago. Thanks for giving me hope!

just_another_person@lemmy.world · 4 months ago

deleted by creator

Nickm8@lemmy.world · 4 months ago

Have been using it a while now, I recommend using something like Tailscale so you can access it from anywhere on your phone. I also have a raspberry pi that can wake up my main machine when I need it.

ooli@lemmy.world · 4 months ago

where the link to the download?

The Hobbyist@lemmy.zip · 4 months ago

whats great is that with ollama and webui, you can as easily run it all on one computer locally using the open-webui pip package or in a remote server using the container version of open-webui.

Ive run both and the webui is really well done. It offers a number of advanced options, like the system prompt but also memory features, documents for RAG and even a built in python ide for when you want to execute python functions. You can even enable web browsing for your model.

I’m personally very pleased with open-webui and ollama and they both work wonders together. Hoghly recommend it! And the latest llama3.1 (in 8 and 70B variants) and llama3.2 (in 1 and 3B variants) work very well, even on CPU only, for the latter! Give it a shot, it is so easy to set up :)

Tobberone@lemm.ee · 4 months ago

Do you know of any nifty resources on how to create RAGs using ollama/webui? (Or even fine-tuning?). I’ve tried to set it up, but the documents provided doesn’t seem to be analysed properly.

I’m trying to get the LLM into reading/summarising a certain type of (wordy) files, and it seems the query prompt is limited to about 6k characters.

jonno@discuss.tchncs.de · 4 months ago

Are you running these llms in containers completely cut off from the internet? My understanding was that the “local first” llms aren’t truly offline and only try and answer base queries offline before contacting their provider for support. This invalidating the privacy argument.

The Hobbyist@lemmy.zip · 4 months ago

The interface called open-webui can run in a container, but ollama runs as a service on your system, from my understanding.

The models are local and only answer queries by default. It all happens on the system without any additional tools. Now, if you want to give them internet access, you can, it is an option you have to setup and open-webui makes that possible though I have not tried it myself. I just see it.

I have never heard of any llm “answer base queries offline before contacting their provider for support”. It’s almost impossible for the LLM to do it by itself without you setting things up for it that way.