Israel bombs Iranian state TV during live broadcast

brucethemoose@lemmy.world · 1 hour ago

Eh, there’s not as much attention paid to them working across hardware because AMD prices their hardware uncompetitively (hence devs don’t test them much), and AMD themself focuses on the MI300X and above.

Also, I’m not sure what layer one needs to get ROCM working.

brucethemoose@lemmy.world · edit-2 2 hours ago

The pinned Tweet on Alex’s Twitter page:

https://xcancel.com/realalexjones

Wednesday LIVE: Desperate Deep State/MSM Pushes “Epstein Hoax” & “MAGA Divorce” In Attempt To Fracture & Bring Down The Trump Administration! Tune In NOW For Latest Developments As Alex Jones Covers This & Other Key Issues!

Be wary of Rawstory headlines, assuming that’s the one you’re talking about.

brucethemoose@lemmy.world · edit-2 3 hours ago

Alex Jones predates Trump and has always been combative, even about Trump policy, from what I’ve seen. He criticized some of Trump’s earlier moves, like with Ukraine.

This is all in character for him, though I admit that’s just my surface impression.

But objectively, his pinned tweet outright says “MAGA Divorce is a false liberal meme.” That is not the message of someone “disavowing” Trump.

You can try to read between the lines, but IMO you should take what he says at face value.

brucethemoose@lemmy.world · edit-2 3 hours ago

The pinned post is literally:

Wednesday LIVE: Desperate Deep State/MSM Pushes “Epstein Hoax” & “MAGA Divorce” In Attempt To Fracture & Bring Down The Trump Administration! Tune In NOW For Latest Developments As Alex Jones Covers This & Other Key Issues!

Scrolling down, I see clips of him being… combatively apologist. The tone being “Trump is an idiot for doing this, instead of that.”

Also, thank the heavens xcancel is still running from whatever black magic backend they’ve figure out.

brucethemoose@lemmy.world · 4 hours ago

Even the small local AI niche hates ChatGPT, heh.

brucethemoose@lemmy.world · edit-2 4 hours ago

I’m sorry, but this is Rawstory writing what people wanna hear for clickbait. Here’s their clip:

https://rumble.com/v6w9hfk-thats-a-cult-alex-jones-disavows-trump-after-he-tells-fans-to-drop-epstein.html

That is not repudiation. It’s wobbling on a single tweet, even with their selective cut.

Honestly I wish this source was softbanned from /c/news. Like it is from Wikipedia: https://en.wikipedia.org/wiki/Wikipedia:Reliable_sources/Perennial_sources#The_Raw_Story

brucethemoose@lemmy.world · edit-2 6 hours ago

Nope.

They are turning on his underlings, but I have not seen a single major influencer call Trump out directly. “It’s Bondi’s fault,” is the party line, which is unreal given how unflinchingly loyal she’s been.

The headlines claiming Trumps supporters are turning on him feel like clickbait.

brucethemoose@lemmy.world · 7 hours ago

Yeah.

I saw Fox comedy talk like that when I was young, I guess. I only vaguely remember the hosts, but recall some grounding others like that.

brucethemoose@lemmy.world · 8 hours ago

Clickbait.

brucethemoose@lemmy.world · edit-2 8 hours ago

There may be thought in a sense.

A analogy might be a static biological “brain” custom grown to predict a list of possible next words in a block of text. It’s thinking, sorta. Maybe it could acknowledge itself in a mirror. That doesn’t mean it’s self aware, though: It’s an unchanging organ.

And if one wants to go down the rabbit hole of “well there are different types of sentience, lines blur,” yada yada, with the end point of that being to treat things like they are…

All ML models are static tools.

For now.

brucethemoose@lemmy.world · 8 hours ago

brucethemoose@lemmy.world · edit-2 8 hours ago

Because third panel wants to live in their house not take over their country

Again. I will remember this when the US more seriously talks up ‘taking over’ Canada; but not living in their house. Not a crisis of imperialist aggression, by that argument, unless you wish to clarify?

brucethemoose@lemmy.world · edit-2 1 day ago

…Okay.

First of, of course I would. What would make you think I wouldn’t? I’m upset with how many civilians are dying now. And I’m upset with genocides elsewhere.

Second. How does that stop this meme from applying to Russia/Ukraine?

brucethemoose@lemmy.world · edit-2 1 day ago

Nor that Russia was doing settler colonialism.

Ah. It’s just a good old annexation. That’s totally fine.

I’ll remember that when my stupid country makes an attempt to annex Canada… Which is not invading their home, because it’s not settler colonialism?

brucethemoose@lemmy.world · 1 day ago

Me: “Ah, another Ukraine meme…”

lemmy.ml

Me: “Ah. My mistake.” Tips hat

brucethemoose@lemmy.world · edit-2 2 days ago

It depends!

Exllamav2 was pretty fast on AMD, exllamav3 is getting support soon. Vllm is also fast AMD. But its not easy to setup; you basically have to be a Python dev on linux and wrestle with pip. Or get lucky with docker.

Base llama.cpp is fine, as are forks like kobold.cpp rocm. This is more doable without so much hastle.

The AMD framework desktop is a pretty good machine for large MoE models. The 7900 XTX is the next best hardware, but unfortunately AMD is not really interested in competing with Nvidia in terms of high VRAM offerings :'/. They don’t want money I guess.

And there are… quirks, depending on the model.

I dunno about Intel Arc these days, but AFAIK you are stuck with their docker container or llama.cpp. And again, they don’t offer a lot of VRAM for the $ either.

NPUs are mostly a nothingburger so far, only good for tiny models.

Llama.cpp Vulkan (for use on anything) is improving but still behind in terms of support.

A lot of people do offload MoE models to Threadripper or EPYC CPUs, via ik_llama.cpp, transformers or some Chinese frameworks. That’s the homelab way to run big models like Qwen 235B or deepseek these days. An Nvidia GPU is still standard, but you can use a 3090 or 4090 and put more of the money in the CPU platform.

You wont find a good comparison because it literally changes by the minute. AMD updates ROCM? Better! Oh, but something broke in llama.cpp! Now its fixed an optimized 4 days later! Oh, architecture change, not it doesn’t work again. And look, exl3 support!

You can literally bench it in a day and have the results be obsolete the next, pretty often.

brucethemoose@lemmy.world · edit-2 2 days ago

Depends. You’re in luck, as someone made a DWQ (which is the most optimal way to run it on Macs, and should work in LM Studio): https://huggingface.co/mlx-community/Kimi-Dev-72B-4bit-DWQ/tree/main

It’s chonky though. The weights alone are like 40GB, so assume 50GB of VRAM allocation for some context. I’m not sure what Macs that equates to… 96GB? Can the 64GB can allocate enough?

Otherwise, the requirement is basically a 5090. You can stuff it into 32GB as an exl3.

Note that it is going to be slow on Macs, being a dense 72B model.

brucethemoose@lemmy.world · edit-2 2 days ago

One last thing: I’ve heard mixed things about 235B, hence there might be a smaller, more optimal LLM for whatever you do.

For instance, Kimi 72B is quite a good coding model: https://huggingface.co/moonshotai/Kimi-Dev-72B

It might fit in vllm (as an AWQ) with 2x 4090s. It and would easily fit in TabbyAPI as an exl3: https://huggingface.co/ArtusDev/moonshotai_Kimi-Dev-72B-EXL3/tree/4.25bpw_H6

As another example, I personally use Nvidia Nemotron models for STEM stuff (other than coding). They rock at that, specifically, and are weaker elsewhere.

brucethemoose@lemmy.world · edit-2 2 days ago

Ah, here we go:

https://huggingface.co/ubergarm/Qwen3-235B-A22B-GGUF

Ubergarm is great. See this part in particular: https://huggingface.co/ubergarm/Qwen3-235B-A22B-GGUF#quick-start

You will need to modify the syntax for 2x GPUs. I’d recommend starting f16/f16 K/V cache with 32K (to see if that’s acceptable, as then theres no dequantization compute overhead), and try not go lower than q8_0/q5_1 (as the V is more amenable to quantization).

brucethemoose@lemmy.world · edit-2 2 days ago

Qwen3-235B-A22B-FP8

Good! An MoE.

Ideally its maxium context lenght of 131K but i’m willing to compromise.

I can tell you from experience all Qwen models are terrible past 32K. What’s more, going over 32K, you have to run them in a special “mode” (YaRN) that degrades performance under 32K. This is particularly bad in vllm, as it does not support dynamic YaRN scaling.

Also, you lose a lot of quality with FP8/AWQ quantization unless it’s native FP8 (like deepseek). Exllama and ik_llama.cpp quants are much higher quality, and their low batch performance is still quite good. Also, VLLM has no good K/V cache quantization (its FP8 destroys quality), while llama.cpp’s is good, and exllama’s is excellent, making it less than ideal for >16K. Its niche is more highly parallel, low context size serving.

My current setup is already: Xeon w7-3465X 128gb DDR5 2x 4090

Honestly, you should be set now. I can get 16+ t/s with high context Hunyuan 70B (which is 13B active) on a 7800 CPU/3090 GPU system with ik_llama.cpp. That rig (8 channel DDR5, and plenty of it, vs my 2 channels) should at least double that with 235B, with the right quantization, and you could speed it up by throwing in 2 more 4090s. The project is explicitly optimized for your exact rig, basically :)

It is poorly documented through. The general strategy is to keep the “core” of the LLM on the GPUs while offloading the less compute intense experts to RAM, and it takes some tinkering. There’s even a project to try and calculate it automatically:

https://github.com/k-koehler/gguf-tensor-overrider

IK_llama.cpp can also use special GGUFs regular llama.cpp can’t take, for faster inference in less space. I’m not sure if one for 235B is floating around huggingface, I will check.

Side note: I hope you can see why I asked. The web of engine strengths/quirks is extremely complicated, heh, and the answer could be totally different for different models.