• 10 Posts
  • 3K Comments
Joined 2 years ago
cake
Cake day: March 22nd, 2024

help-circle
  • First of all, I mean zero offense with any purchase decision. A 5090 is very good.

    …But if I were paying that kind of money, I’d probably get a 4090 and a new motherboard/CPU instead. Maybe a used DRR4 threadripper system.

    Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.

    Having a fast CPU, with lots of RAM channels, with full PCIe bandwidth is much more important for that than having a 5090, where a 3090 or 4090 will get the job done.

    Even if pure speed is your primary concern, you can tune a sparse 120B model (like Laguna) to run almost as fast as Qwen 27B on a 5090, and get at-least-good results.


    It’s more finicky and involved, though. For sure.

    Running an LLM on a 5090 is a task. Hybrid CPU + GPU inference is a hobby.



    • It’s just under 300B, trained at FP4; absolutely the perfect size for servers with 128GB-192GB CPU RAM to spare.

    • Its fast. I’m getting 17 tokens/sec on a single RTX 3090 GPU, all experts offloaded to RAM; for a 300B model this smart, that’s crazy fast.

    • Its attention mechanism is cutting edge, good for long context without too much processing time.

    • The benchmarks for coding/agenic usage are absolutely bonkers, within margin of error of frontier models or Deepseek Pro in some cases: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

    • …Though I don’t put much stock in benchmarks. And I haven’t tested it enough to tell you if it lives up to that hype for specific use cases.

    • Deepseek also publishes its base model. That means I can “unfry” the model with a merge if I have to.


    I am afraid of the the model being “overfit” to coding and agenic stuff.

    For reference, my previous favorite model was Xiaomi MiMo V2.5 310B. It benched well, but it also feels “smart” outside of benchmarks, like in knowledge of trivia without tool usage/internet access or comprehension of weird questions.



  • [AIT] I know this isn’t everyone’s cup of tea, but I’m excited about Deepseek V4 flash. It’s (for me) the perfect size and architecture to self-host an LLM.

    My box (and brain) are chugging through a queue:

    • Figure out why my swap is going crazy, and how to ban processes from it [Done].

    • Figure out why Code OSS is unhappy [Partially Done].

    • Make an ik_llama.cpp iMatrix for Deepseek V4 [Done].

    • Figure out why quantization isn’t working [Done].

    • Make a test IQ2_KL/MXFP4_R8 quant to see how it does squeezed onto my box [in progress].

    • Test. Tune. Inevitably troubleshoot the dozen other things that go wrong. Figure out how much spare RAM that leaves me.

    • Make a higher quality IQ3_KT quantization. This will take all night on my CPU.

    • KLD test both of them vs the full precision, to quantify quantization loss. Likely an overnight task, too.

    • Try merging the new model release with the base model, 50/50, for a less “deep fried” model. imatrix, quant, test.

    The goal is to host it on a single RTX 3090, Ryzen 7000 with 128GB RAM, for anyone curious. Though I may try smaller models too, like Laguna S1.


  • It’s just a product of optimizing for engagement maximization at all costs, like tons of other products from Big Tech.

    Much of the sycophancy and love-bombing you see in modern LLMs comes from RLHF:

    https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback

    They’re engineered as technical tools, and outputs humans prefer are selected in training.

    …The problem is that Big Tech wrapped what’s really glorified autocomplete or a coding/text processing bot, optimized to please humans with its output, in extensive scaffolding to make it look like an omniscient oracle. Every word that comes out of Altman’s mouth is textbook cult language, and they engineer that “vibe” into the system prompt and RAG system.


    Take the raw LLM weights by themself, run in a purely technical context… and it’s obviously not a god. Even to the average person, its more like fancy spellcheck; a tool.

    It’s all the dressing-up, hype, deliberate obfuscation, the positive feedback loops and such that turn them into a machine that can perpetuate a cult.


    What I’m trying to say is it’s not “runaway technology.”

    It’s people.

    It’s people doing this on purpose.

    “Blaming” the LLM weights instead dilutes their responsibility. Every ounce of blame is on people like Altman or Modi, but THEY want you to think it’s out of their control, hence the whole quasi religious angle.










  • brucethemoose@lemmy.worldtoFediverse@lemmy.worldSome thoughts after 3 years on Lemmy
    link
    fedilink
    English
    arrow-up
    52
    arrow-down
    4
    ·
    edit-2
    3 days ago

    Also, to add:

    Information hygiene here is bad.

    Every day, nearly every time I scroll the feed, some spicy headline goes past. When I go to look it up, it’s often completely made-up misinformation, or a warped ragebait headline, or a literally bot-written Tweet, yet it’s voted right to the front page.

    …I thought Reddit refugees would be sick of this, and more careful/skeptical, but apparently not. I’ve even talked to some users who don’t understand the distinction between opinion columns and news.

    And I had an argument with a mod, who had no problem with an LLM engagement farm Tweet as long as its “the right message.”


  • Maybe Debian could create its own LLM trained on Debian for Debian? cuts out the licensing issues and keeps the community in control?

    That’s… not how it works, unfortunately.

    I think what you’re getting at would be strict criteria along the lines of:

    • A human must take full responsibility for any contribution.
    • And only with fully disclosed, labeled assistance from “legally unproblematic” LLMs. The Nemotron or Olmo series would be examples: apache licensed, open and auditable datasets, can be fully controlled by Debian community.

    They could enforce some specific harness, or mathematically verifiable output, like some other projects already do. I know less about the technical details of that, though.


  • brucethemoose@lemmy.worldtoFediverse@lemmy.worldSome thoughts after 3 years on Lemmy
    link
    fedilink
    English
    arrow-up
    21
    arrow-down
    32
    ·
    edit-2
    3 days ago

    On the contrary, I think the impulse to check a user’s post history and judge them is a bad habit on Lemmy.

    Take me. Scroll through my history, and you’ll probably peg me as a AI Tech Bro and a… closet neolib, maybe? I dunno. But if it’s along those lines, you could not be more wrong, on either count.

    With the exception of extreme trolls (or moderators doing their thing), I think Lemmy’s culture of “purity checks” is poison to the platform. It turns communities insular, toxic, and angry like it does on Reddit, governed by takes that have nothing to do with the particular thread. It kills the focus of communities. It rewards conformance.

    And this thread is a good example. I thought OP was raising some good general points, but digging previous comments up threw all that out the window. I’ve seen this happen in other threads, where a technical or fandom discussion devolves into unrelated, raging arguments on something I did not come there to read.


  • This reflects my experience, approximately.

    Yeah… Lemmy/Piefed is a mixed experience for niche interests. I fear a set of “insulated hyper-political echo chambers,” to put it somewhat crudely, is not going to attract the masses from corporate social media.


    It’s kinda my fault, though. I’m not putting enough effort in seeking out those niches, much less posting in them; it’s too easy to doomscroll.

    I keep telling myself to dedicate more attention to “specialized” forums like HN or SpaceBattles or DPReview, but alas, my brain keeps failing me there too.


  • That’s not even “breakthroughs,” though.

    That’s just “follow instructions for what’s already been done, and meet par.” Anyone with some Python experience and instruction following ability could do it. In LLM land, what some research houses are doing would be like releasing some new white box computer with chips from 15 years ago.

    It makes no sense. It’s way too late to be interesting.

    But I guess it’s getting funded because the funders can’t really tell.


  • Many LLM labs are shockingly dysfunctional.

    Amazon and the useless Nova series is a great example, but Microsoft has its own too. AMD as well.

    Europe has a ton of them, labs barely replicating obsolete Llama2 training routines in 2026, but somehow getting tons of taxpayer dollars? Saudi labs doing basically nothing.

    I mean… I could follow an off-the-shelf Nvidia Nemotron recipe and do better. Or more sanely, just start with that as a base. And that’s sad; apparently these highly paid labs don’t even follow the space enough to know to do that? What on Earth are they even doing internally?