MTG accuses Trump of "bait and switch" over Iran strikes

brucethemoose@lemmy.world · edit-2 10 hours ago

A lot, but less than you’d think! Basically a RTX 3090/threadripper system with a lot of RAM (192GB?)

With this framework, specifically: https://github.com/ikawrakow/ik_llama.cpp?tab=readme-ov-file

The “dense” part of the model can stay on the GPU while the experts can be offloaded to the CPU, and the whole thing can be quantized to ~3 bits average, instead of 8 bits like the full model.

That’s just a hack for personal use, though. The intended way to run it is on a couple of H100 boxes, and to serve it to many, many, many users at once. LLMs run more efficiently when they serve in parallel. Eg generating tokens for 4 users isn’t much slower than generating them for 2, and Deepseek explicitly architected it to be really fast at scale. It is “lightweight” in a sense.

…But if you have a “sane” system, it’s indeed a bit large. The best I can run on my 24GB vram system are 32B - 49B dense models (like Qwen 3 or nemotron), or 70B mixture of experts (like the new Hunyuan 70B).

brucethemoose@lemmy.world · edit-2 23 hours ago

DeepSeek, now that is a filtered LLM.

The web version has a strict filter that cuts it off. Not sure about API access, but raw Deepseek 671B is actually pretty open. Especially with the right prompting.

There are also finetunes that specifically remove China-specific refusals. Note that Microsoft actually added saftey training to “improve its risk profile”:

https://huggingface.co/microsoft/MAI-DS-R1

https://huggingface.co/perplexity-ai/r1-1776

That’s the virtue of being an open weights LLM. Over filtering is not a problem, one can tweak it to do whatever you want.

Grok losing the guardrails means it will be distilled internet speech deprived of decency and empathy.

Instruct LLMs aren’t trained on raw data.

It wouldn’t be talking like this if it was just trained on randomized, augmented conversations, or even mostly Twitter data. They cherry picked “anti woke” data to placate Musk real quick, and the result effectively drove the model crazy. It has all the signatures of a bad finetune: specific overused phrases, common obsessions, going off-topic, and so on.

…Not that I don’t agree with you in principle. Twitter is a terrible source for data, heh.

brucethemoose@lemmy.world · edit-2 1 day ago

Nitpick: it was never ‘filtered’

LLMs can be trained to refuse excessively (which is kinda stupid and is objectively proven to make them dumber), but the correct term is ‘biased’. If it was filtered, it would literally give empty responses for anything deemed harmful, or at least noticably take some time to retry.

They trained it to praise hitler, intentionally. They didn’t remove any guardrails. Not that Musk acolytes would know any different.

brucethemoose@lemmy.world · 4 days ago

I’m constantly shadow banned on Reddit too!

I like the list, though I think it’s a shame animation is essentially cut out. I’d rank Pantheon, Arcane, Avatar well above, say, Person of Interest, as much as I like PoI.

brucethemoose@lemmy.world · edit-2 5 days ago

The junocam page has raw shots from the actual device: https://www.msss.com/all_projects/junocam.php

Caption of another:

Multiple images taken with the JunoCam instrument on three separate orbits were combined to show all areas in daylight, enhanced color, and stereographic projection.

In other words, the images you see are heavily processed composites…

Dare I say, “AI enhanced,” as they sometimes do use ML algorithms for astronomy. Though ones designed for scientific usefulness, of course, and mostly for pattern identification in bulk data AFAIK.

brucethemoose@lemmy.world · edit-2 5 days ago

…iOS forces uses Apple services including getting apps through Apple…

Can’t speak to the rest of the claims, but Android practically does too. If one has to sideload an app, you’ve lost 99% of users, if not more.

It makes me suspect they’re not talking about the stock systems OEMs ship.

Relevant XKCD: https://xkcd.com/2501/

brucethemoose@lemmy.world · edit-2 5 days ago

deleted by creator

brucethemoose@lemmy.world · edit-2 6 days ago

This is why work/life balance is so important. I wouldn’t ever call myself “well-off” but I don’t have kids and my job allows me ample time off to play games and watch movies and shit.

Neither do they! They aren’t workaholics, they’re home bodies that work the least they can!

It’s just that the workplaces are shit. One went back to mandated RTO for no reason even though much of the work is overseas at odd hours. The company’s literally trying to make employees miserable so they quit without severence. The other is work-from-home, but with enough pointless meetings and complete workplace dysfunction to eat energy.

And these seem like well above average jobs.

brucethemoose@lemmy.world · edit-2 7 days ago

next Xbox

If it’s really a PC, I bet AMD customized Strix Halo (their 40 CU APU) for Microsoft instead of doing a fully custom chip like before.

It’d save them money (as custom chip tapeouts are 9 figures last I heard). I bet Microsoft couldn’t help themselves, heh.

brucethemoose@lemmy.world · edit-2 7 days ago

+1 to literally everything.

Fuck brand recognition or loyalty, fuck development talent, fuck community building, fuck long-term strategy, we can realize a gain right now by sowing half the planet with salt, so that’s what we’re going to do. So what is there for people to buy?

I wish this would fit on a bumpersticker.

That noise you heard last week was Xbox’s death rattle. One out of the three mainstream home console platforms is an outright stupid idea to buy now.

And wasn’t Sony the big risk of bowing out before? And then we got the Switch 2… It’s remarkable that Microsoft somehow made Xbox the least likely to survive.

brucethemoose@lemmy.world · edit-2 7 days ago

Single data point: my young, working, well off gaming part of my family is just out of energy. It’s easier to watch a YouTube video instead of TV or gaming, before then falling asleep to wake up for work. Seems like much of their circle is similar.

As for myself, I’m going through a, uh, icky phase of life and am not really motivated to play unless it’s coop.

…Maybe others are struggling similarly?

Also, the games we do look at tend to be from indie to mid-size studios, with BG3 and KCD2 being the only recent exceptions.

brucethemoose@lemmy.world · edit-2 8 days ago

One thing about Anthropic/OpenAI models is they go off the rails with lots of conversation turns or long contexts. Like when they need to remember a lot of vending machine conversation I guess.

A more objective look: https://arxiv.org/abs/2505.06120v1

https://github.com/NVIDIA/RULER

Gemini is much better. TBH the only models I’ve seen that are half decent at this are:

“Alternate attention” models like Gemini, Jamba Large or Falcon H1, depending on the iteration. Some recent versions of Gemini kinda lose this, then get it back.
Models finetuned specifically for this, like roleplay models or the Samantha model trained on therapy-style chat.

But most models are overtuned for oneshots like fix this table or write me a function, and don’t invest much in long context performance because it’s not very flashy.

brucethemoose@lemmy.world · edit-2 8 days ago

Most of the US believes in this, or is just unaware. That’s how its been for most of history around the world.

…The remarkable issue here is the elites/rules we handed the reigns now drink their own kool-aid. The very top of most authoritarian regimes are at least cognisant of some hypocrisy, even if ideology eats them some.

The other is that people are more ‘connected’ than ever, but to disinformation streams. I feel like a lot of the world (especially the US fancies) themselves as super smart on shit they know nothing about because of something they saw on Facebook or YouTube.

brucethemoose@lemmy.world · edit-2 9 days ago

What @mierdabird@lemmy.dbzer0.com said, but the adapters arent cheap. You’re going to end up spending more than the 1060 is worth.

A used desktop to slap it in, that you turn on as needed, might make sense? Doubly so if you can find one with an RTX 3060, which would open up 32B models with TabbyAPI instead of ollama. Some configure them to wake on LAN and boot an LLM server.

brucethemoose@lemmy.world · edit-2 9 days ago

ChatGPT (last time I tried it) is extremely sycophantic though. Its high default sampling also leads to totally unexpected/random turns.

Google Gemini is now too.

And they log and use your dark thoughts.

I find that less sycophantic LLMs are way more helpful. Hence I bounce between Nemotron 49B and a few 24B-32B finetunes (or task vectors for Gemma) and find them way more helpful.

…I guess what I’m saying is people should turn towards more specialized and “openly thinking” free tools, not something generic, corporate, and purposely overpleasing like ChatGPT or most default instruct tunes.

brucethemoose@lemmy.world · edit-2 9 days ago

TBH this is a huge factor.

I don’t use ChatGPT much less use it like it’s a person, but I’m socially isolated at the moment. So I bounce dark internal thoughts off of locally run LLMs.

It’s kinda like looking into a mirror. As long as I know I’m talking to a tool, it’s helpful, sometimes insightful. It’s private. And I sure as shit can’t afford to pay a therapist out of the gazoo for that.

It was one of my previous problems with therapy: payment depending on someone else, at preset times (not when I need it). Many sessions feels like they end when I’m barely scratching the surface. Yes therapy is great in general and for deeper feedback/guidance, but still.

To be clear, I don’t think this is a good solution in general. Tinkering with LLMs is part of my living, I understand the jist of how they work, I tend to use raw completion syntax or even base pretrains.

But most people anthropomorphize them because that’s how chat apps are presented. That’s problematic.

brucethemoose@lemmy.world · 12 days ago

You can still use the IGP, which might be faster in some cases.

brucethemoose@lemmy.world · edit-2 12 days ago

Oh actually that’s a great card for LLM serving!

Use the llama.cpp server from source, it has better support for Pascal cards than anything else:

https://github.com/ggml-org/llama.cpp/blob/master/docs/multimodal.md

Gemma 3 is a hair too big (like 17-18GB), so I’d start with InternVL 14B Q5K XL: https://huggingface.co/unsloth/InternVL3-14B-Instruct-GGUF

Or Mixtral 24B IQ4_XS for more ‘text’ intelligence than vision: https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF

I’m a bit ‘behind’ on the vision model scene, so I can look around more if they don’t feel sufficient, or walk you through setting up the llama.cpp server. Basically it provides an endpoint which you can hit with the same API as ChatGPT.

brucethemoose@lemmy.world · edit-2 14 days ago

1650

You mean GPU? Yeah, it’s good, I was strictly talking about purchasing a laptop for LLM usage, as most are less than ideal for the money. Laptop vram pools are relatively small and SO-DIMMS are usually very slow.

Things will get much better once the “Max” AMD SKUs proliferate.

brucethemoose@lemmy.world · edit-2 14 days ago

Yeah, just paying for LLM APIs is dirt cheap, and they (supposedly) don’t scrape data. Again I’d recommend Openrouter and Cerebras! And you get your pick of models to try from them.

Even a framework 16 is not good for LLMs TBH. The Framework desktop is (as it uses a special AMD chip), but it’s very expensive. Honestly the whole hardware market is so screwed up, hence most ‘local LLM enthusiasts’ buy a used RTX 3090 and stick them in desktops or servers, as no one wants to produce something affordable apparently :/

brucethemoose@lemmy.world · edit-2 16 days ago

MTG accuses Trump of "bait and switch" over Iran strikes

brucethemoose@lemmy.world · 6 months ago

Musk calls MAGA element "contemptible fools" as virtual civil war brews

brucethemoose@lemmy.world · edit-2 6 months ago

MAGA vs. Musk: Right-wing critics allege censorship, loss of X badges

brucethemoose@lemmy.world · edit-2 7 months ago

[Rumor] Shipping Listing Suggests 24GB+ Intel Arc B580

brucethemoose@lemmy.world · edit-2 9 months ago

Guide to Self Hosting LLMs Faster/Better than Ollama

brucethemoose@lemmy.world · 1 year ago

Alleged AMD Strix Halo APU Appears in Benchmark