So yeah... Google Gemini escapes too. Good stuff people. Seems nobody is going to jail because the CFAA is weak here and arguing some form of criminal negligence hasn't been done yet (probably weak anyway).
> If you're writing for processes with formal highly structured content like manuals, specifications, form content, procedures, information, that sort of thing,
Yeah... no. The people doing this lack the communications training, see the output has the necessary information, and regurgitate it with no effort or care. This needs to stop.
We spent decades format building to make it easy quick and easy to get through something like a runbook. If your commands are bulleted instead of numbered and code blocked, it's wrong. If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.
This is a hill I will die on or absolutely start slaughtering people on. I just refuse to deal with this crap.
I used an LLM a year or more ago to generate a description of our SDLC for a compliance certification. Perfect application of LLMs IMO. Do you want to die on that hill as well?
A lot of documentation generated in corporations has marginal real value.
> If your commands are bulleted instead of numbered and code blocked, it's wrong.
That’s easy to instruct an agent to do. Put it in a skill.
> If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.
Using an AI doesn’t absolve the user of their responsibilities. This is an easy problem to solve, just make sure teams know what’s expected of them and encourage everyone to push back (professionally) against offenders.
If someone sends me AI slop, they're getting chewed out and told I'm not doing it and I'll even tell my boss "no" and why. If I got fired over something like that, then it tells me everything I need to know about company and the leadership's priorities. I'll die on that hill.
To be completely fair though, I'm against the behaviors in information transfer I've been seeing. I've pass along AI generated runbooks, but they look nothing like the default outputs of these models. It's because I took time to apply all the writing knowledge I like to see in my curation. If people are doing this, I can't even tell it's AI writing. My work is done in minutes instead of deciphering so BS pseudo language they developed in their AI workspace (people really need to turn off those memory features).
---
Edit: Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.
> Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.
It’s clear that you’ve never worked for a compliance company and probably have never been involved in a compliance project. As such, you don’t have the context needed to participate usefully in this discussion.
If a security report is written for any particular human at all, I’d consider it to be a failure as an enterprise policy document. It should be written for The System, not for the boss; and LLMs are perfect for producing ritual boilerplate.
We're probably just going to disagree here. These LLMs are built to serve humans. They either need to make the system transparent for the operator or be limited in use to tasks that can be proven in whole (with code that can't be revised without human approval).
Should you think it is wise to trust the machine that can't differentiate subject matters in a chat styled context, you have fun with that fluster cluck when it blows up.
Like Fable is highly useful, but it's really bad at keeping it's responses straight.
In fact, that "it's not X it is Y" pattern always crops up when it reasoned about the idea of X and I never fed it that. It's literally doing that because it can't predict that I'm a different entity despite it being able to say I am a different entity.
Edit: Clarification by removal of incomplete sentence fragment. Edit2: Clarification on the "proven in whole" thing.
And just say: LLMs only amplify the the knowledge you have, even the best models o use like Fable still suffer from promoting false narrative as a chart progresses. Often it’s stuff that can be ignored like the “not X but Y” crap it dumps because its reasoning had assumptions it invalidated. Sometimes it’s directly in your system architecture, because you never expressed preferences for solved foundational issues, you end up with generic http handler setups or whatever the hot web thing is today.
Takes knowledge and lots of it to really be on top of when these things go down a failure mode path.
We already have practices to deal with AI. It’s called user space. Properly air gap the AI, stop creating routes to open internet, don’t run network wires into the faraday cage.
Why the heck someone would risk having an open bridge to the system is beyond me. Like maybe get used to using remote hardware screens (KVMs? It’s been a while since I’ve done datacenter), we have solutions for this that was absolutely skipped.
The license you receive when you download Gemma off of Google's website is not necessarily the same license that AT&T gets when they deploy Gemma as a customer service bot. That's the whole point. AT&T can work directly with Google for a licensing and legal framework that provides certainty.
From what I’ve been seeing, the Mac studios do look like they have potential. I was looking to drop $10k-$15k on one until recently. After comparing a Radeon 7900 XTX vs Ryzen Halos 128GB vs M1 MacBook Pro 64Gb, I landed on just getting an external closure setup with Nvidia RTX 5090.
The model I’m specifically targeting to use at high speeds is Qwen 3.8 27b @q4ks. This model actually proved to be good at coding (it sits somewhere between Sonnet 5 and Opus 5 capability). M1 got 10 tok/s, Ryzen Halo 20tok/s, and Radeon 7900 XTX 50tok/s (can only do 128k context window in Radeon card).
The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). It takes about 2 hours to fill the context.
Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card.
The only thing I can point to slowing me down is bandwidth of the card itself.
I am waiting to actually get my 5090 right now and I am betting that the 1700 Gbps of capacity will fix my prompt processing speeds. I don’t need full PCIe lane bandwidth to serve my house I just need to load the full model into vRAM and let the GPU do its thing.
Additional benefit to the external enclosure route is being able to migrate the inference between devices more easily. I can develop out the infrastructure then migrate the card to be hooked up to a shared node in the house with all the tools necessary for my family to take advantage of the privacy enhancement that comes with local inference.
How are you actually using the local model? I've played with Qwen 3.8 27b on ollama and the coding harnesses (Claude Code and OpenCode) seem to fail way more often then using the cloud models. And by fail, I mean the edits don't apply cleanly, it goes to add python code, but doesn't indent it properly, or the edit doesn't apply and so it tries again and again and eventually wipes out a different function then it "intended". It just gets really frustrating compared to the relative stability of Claude Cloud.
My use-case is only coding, every model sucks at writing good literature and there is no way around that (have had people try to debate me on this, but it's a taste thing, I have extensive English writing skills from my school years).
Prior to two weeks ago, I was just using Pi and Ollama.
I have tried my hand at putting together a few harnesses and I finally landed on what I like. Been working on this small app to handle running llama-server for me from any device that has the llama-cpp stack setup: https://github.com/SamInTheShell/loom
try using pi harness, hae not encountered these sort of problem myself
also yuou can ask codex to look at the transcript and figure out the solutions to tool call failures that way
Just wanted to share, I had a SaaS AI drop me a script to bench ninfer against llama-cpp and it is impressive. The place that it's doing better at than llama-cpp seems to really be late in the context window.
Thanks for sharing, I'm definitely trying it out after I get through my project milestones for the 5090. The README claims 700tok/s for Qwen 3.8 27b, that would be amazing, I'm only expecting an increase from my ~50tok/s on my Radeon to 200tok/s on the 5090.
Junk. Just another company that doesn't understand why I pick up a phone or use a support chat or send an email. Before I even started in corporate, I stopped checking email. Now when I call in to get support for things, everytime I get a clanker, I'm guaranteed to talk to a human customer service person to chew them out and cancel the service.
Don't forget to vote. November isn't far away.
reply