Hacker Newsnew | past | comments | ask | show | jobs | submit | CharlieDigital's commentslogin

Interesting.

I wanted to test the veracity of this result so I tested some of my own 100% hand written pieces and what do you know? It came back 100% hand written.

This is a better ad for Pangram than for OP :D


- This is only true if your constraint is online

- If it is offline processing (e.g. classifying text snippets stored in a DB (from call transcript, from logs)), you can batch to Luna

- If you batch to Luna and then divide both cost and time by the batch size, you will find that Luna beats Jev.

- Now tune your batch size for your test dataset and see where the batch size causes accuracy falloff.

Tested this approach with GPT-6 Luna. It was 1.6x cheaper, 1.2x slower, within margin of error performance vs Jev with batch size 20 using a CFPB complaint dataset. Tuning batch size up to even double would likely yield similar accuracy while reducing both the per-record run time as well as per-record token cost.

The methodology of comparing single record is only valid for the on-line, real-time use case. For every other case, Luna can match or beat Jev by simply batching. If you find no dropoff at larger batch sizes, you will be significantly cheaper than Jev.

* Batching here does not mean the native batch API but actually placing 20 records (batch size) into one prompt and getting 20 results back in one response.


    > The model matters more than the harness anyway
    > 
    > Everyone is benchmaxxing
    > 
    > ...harnesses tend to be chosen on voodoo and hunches...
I get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost. The whole point of their technical implementation and design decision here is to highlight that it's not "voodoo and hunches", but observable data.

Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less. So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.

The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on "voodoo and hunches" by using actual data to back the assertions. Your comment feels misguided and completely hand waves the actual data points here.


It won't be Pi because there's really no singular Pi that is broadly useful without explicit configuration of plugins. Pi requires plugins to do lots of things that are OOB on other harnesses that people care about. MCP, OpenTelemetry, etc. It may be some offshoot or something built on top of Pi that is more standardized, but it won't be Pi.

Yes, definitely fair. Ideally all the models would be trained against "Pi in exactly the configuration I personally prefer", but alas :)

Company's own website job listings are likely actually in an applicant tracking system (ATS) like Greenhouse or Ashby because they need to manage the pipeline, not just list the job.

You pay for job listings so it would be silly to not take them down.

I can already see it. 7-Nebula, 8-Galactic, 9-Cosmos; The size inflation is real.

It's just another tool. Luna exists for a reason: it's the right tool for the job. If they release AGI and it costs $1 and 5 seconds to decide "is the customer asking for a refund", then that's a terrible use case for AGI if another tool can do it with 95% accuracy for $0.002 and 50ms.

> If they release AGI and it costs $1 and 5 seconds to decide "is the customer asking for a refund", then that's a terrible use case for AGI

Is it? If AGI is here then by the time I test and deploy that the AGI will be most likely cheaper and smarter because it improved itself (for example by implementing it's own Jev for stupid prompts like this), so why invest into a more complex solutions?


I don’t think that was a great example since there’s only so many refunds a customer is going to ask for. And it’s saving time that otherwise maybe would have to go through a human. The rate is low enough that a more expensive model makes more sense.

Though for tasks where you are trying to search through billions of documents, social media posts, etc. and extract certain information, where each individual post is of low value and only the data in aggregate is valuable, then that’s where you’d want something cheaper and faster.

Such as if you want to look at all posts on X in the last few months and find how many have a negative or positive sentiment about the economy (or are unrelated).

Of course you could use a special-purpose model for this, but the whole point of something like Jev is to ask whatever questions you want without having to train something new.


    > so why invest into a more complex solutions
Not sure what's more complex about one REST API call versus another REST API call...

Because AGI will also handle whatever is happening after your "is the customer asking for a refund?" question. Replacing whoever is doing that refund.

Well sure, in the future you may also be able to ask AGI to "please just run my life", while you stay in bed.

In the meantime, today, in the real world, there are businesses wanting to automate well-defined business flows, who don't want some stroppy AGI with a mind of it's own to instead decide to hack into something, or reward hack and make the customer happy by just wire transferring $1M of company money into their account.


Why would it need to? There is a deterministic flow here for the actions that are allowed. AGI isn't needed for this at all if you can map out the flow and use a classifier to decide which route to follow.

    > Fact is, vibe-coded projects devolve over time into an unmaintainable mess. The reason is simple, yet hard to fix: code maintainability and good architecture don’t have good measurements that we can apply, because it takes months, years even, to notice the effects of bad architecture or of unmaintainable code.
    > 
    > For one, AI is not trained on what it means for code to be maintainable. For instance, any reinforcement learning done needs a reward signal that can be measured immediately, not in months or years.
Sad to say, but this is no different from human written code. Human written code just takes even longer to realize the mistakes because the pace is slower.

I think at the end of the day, it is not impossible to have AI write "good" or "high quality" code. If anything, once the patterns are established, AI will be more likely to adhere to the patterns and rules than any human team. It requires the most experienced engineers on the team to split their time writing the core patterns and documenting them in references/skills.

But it takes a lot of "taste" and a willingness to slow down a bit with AI (to create necessary artifacts), something teams find hard to do when you can ship so fast now.

My experience has been that there is a camp of very senior engineers that are unwilling to adapt to reality and focus on documentation and writing (effectively producing skills and agent guidance which multiplies their effectiveness); they will cling to their knowledge thinking coding a sacred art.


> Sad to say, but this is no different from human written code.

I don't think so. It's true that human also write shitty code but the key difference is we actually remember what is the intention behind those crappy implementations so someone can fix it later. aka it is the matter of long term memory that currently LLM architecture is not capable of.

You can argue that claude can read the whole linux codebase and report bugs, but they can only report local bugs, not systematic one. 1M context windows seems like huge, but the effective range is actually pretty limited, and it still does not equal to human insight.


    > aka it is the matter of long term memory that currently LLM architecture is not capable of
Long term memory is easier than you think when you consider what an agent has to do when it is reading and editing code: instruct the agent to leave comments on its rationale and reasoning directly in the code. This is infrastructure free memory that every agent that then sees the code will read. Your code review agent will see the reasoning and decision making your coding agent formulated. When an agent comes and refactors this code in 6 months, the comments will be there (and it will update it!). When an agent is trying to troubleshoot an issue, it will read the comment. No infrastructure needed! Don't overthink it; use comments.

Code comments are line-of-sight for agents and one of the cheapest, highest leverage ways to get better coding performance from AI because unlike skills that may or may not activate, comments end up in context as long as they are well placed and carry the right instructions.

Best places to have it leave comments: 1) start of the file because it frequently uses `sed -n 1,200p` to read files and 2) inside the body of the method because it may find by keyword and read a few lines past. If your harness is set up with an LSP, language standard comments are also useful because then it can read comments on the member.

Tips for comments: point it to other, related members or artifacts; point it to external canonical docs; point is to a specific issue number or PR; have examples directly in the comment using your language's example markers; point it to example, reference usages in code. Use AGENTS.md to tell your agents how you want it to leave comments and to specifically read, follow, and maintain comments.

You don't need infrastructure or special architecture; Every coding agent is text-in, text-out. You need comments that get carried with text-in and a bit of guidance to the agent on how to use comments effectively.


No, developers definitely do not remember what they did two months ago. If you are busy, even two weeks is a problem. That is why we discuss documentation so much, self-documenting code, tickets and tests.

Well, and "intention" is a mine field of its own.


I've seen LLMs "connect the dots" across complex systems many times before. When it works, it's shocking how quickly it can pin down a bug that spans across the software stack.

1M context window is plenty. Once it's skimmed the code and come up with a theory for the problem, it can spin up a subagent that has a whole fresh context window and it can dedicate the whole thing to that one hunch.


To me the difference is humans (ideally) will learn when they build something in a non-optimal way, and so will improve over time to become a competent engineer / architect. We cannot be perfect but to me a huge part of life is learning from failure and improving yourself, something that LLMs short-circuit and cannot replace.

LLMs cannot truly learn and so are destined to produce whatever the "average" software looked like at their training cutoff, or worse to produce code based on _other_ LLM generated code.

Ouroboros eat your heart out


LLMs learn, and in two main ways: in-context and in training stages, release to release. The former is quick and sample efficient - perfect for adjusting AI behavior on the fly, and for enabling AI's own problem-solving capabilities. The latter modifies the "behavior defaults" and gives you performance gains that stick.

Why do you think that "write maintainable code" is somehow impossible to learn for an AI? We already have AI storming the frontiers of research math - way beyond the "average" of the field. If you can RL for "better at math", I see no reason why "better at maintaining code" would be somehow impossible.

You can construct an RL env where a codebase is presented as a "tree", and the AI is given one change to make at a time - and the per-change reward is not just whether the change itself has been evaluated as "made successfully", but also whether it made future changes down the line more or less likely to be successful, and harder or easier to make.

This is a formulation already used by some "maintainable code" benchmarks, so I expect something like it to make is way into frontier lab RL pipelines some time between "next week" and "a couple months ago".


While I mostly agree, I think this is something we need to assume the Pareto principle applies to: likely 20% of humans will improve but 80% will not.

Yes, "code rot" is not in any way an AI-unique problem. Codebases like Flash Player or Bethesda Engine have been deep in decay long before AI was capable of contributing to them.

Historically, this was caused by hiring the cheapest developers one can find, having high turnover, outsourcing, pushing to ship at any cost and more. AI just lets you get there faster, and without having to hire bargain bin Indians.

The thing is, today's AI is already far better at "code rot per feature shipped" than the worst of developers - and I struggle to believe that we're at the limit there.

I've already seen benchmarks that test for AI's ability to make incremental changes and tweaks to code continuously - thus, tracking whether earlier changes make the latter changes harder. This makes for a clear target to RL for.


AI is not better nor worse at producing code rot; just faster at it.

AI produced code is a function of the team driving and instructing the agents along with the scaffolding produced by the team (skills, examples, docs, comments); same with human teams.

A team that cannot guide a human team to produce better code will not be able to guide an AI team to produce better code because it's the same skillset: being able to write good docs, create constraints structurally in code, produce core architecture that enforces good behavior.


Not entirely wrong, but there's a very big hole: "the scaffolding produced by the team" also includes the scaffolding produced by past AIs.

An AI that knows how to keep the documentation accurate and up to date, and does it by default, would, all other things equal, rot your codebase less. An AI that changes the code without checking whether it obsoleted a bunch of examples in the docs would rot your codebase more.

While I think that you can reduce "AI-induced code rot" with good prompting and steering, you could also make headway against it at model level, by making the AI "well-behaved" by default.


    > "the scaffolding produced by the team" also includes the scaffolding produced by past AIs.
This statement is also true of humans. Everything you've stated here is also true for human engineers.

    > "the scaffolding produced by the team" also includes the scaffolding produced by past engineers.
But the agent can be instructed reliably to keep documentation accurate and up to date and will then do so dutifully. Put it in AGENTS.md that it must always update the /docs directory by creating a new doc or updating an existing doc and it will do it. (Yes, adherence may be 95% of the time, but that is likely several points higher than with most non-NASA human teams)

Better yet, extract docs from code comments. Even better when the docs are spatially co-located and line of sight as the agent crawls through code.


> Sad to say, but this is no different from human written code. Human written code just takes even longer to realize the mistakes because the pace is slower.

When the pace is slower you can notice mistakes earlier because you have time to reflect. It also allows you to detect when it’s becoming hard to maintain and you can correct course, rather than after it has become an unworkable mess.


    > because you have time to reflect
It doesn't mean that people do. This is a false narrative we tell ourselves. Yes, there are craft-oriented devs and teams, but these are the exception rather than the rule because in the end, it is the GTM and business teams that define what, when, how and rarely the engineering teams.

There is no team without tech debt because there is no "golden" project where every decision has been made right because of reflection on decisions made wrong.


>It doesn't mean that people do.

After a certain point, people would be forced to refactor, because they find themselves unable to handle the complexity.

With LLMs, there is no such friction. So the complexity get piled upon complexity in the form of a million best practices that is indiscriminately followed...


    > After a certain point, people would be forced to refactor, because they find themselves unable to handle the complexity.
This is a fallacy; this is why legacy code exists that teams just work around. They lack the tests to verify it, the person that wrote it is long gone, it's handling some mission critical dataflow so no one touches the code and just builds around it.

>They lack the tests to verify it, the person that wrote it is long gone, it's handling some mission critical dataflow so no one touches the code and just builds around it.

What you say here is not always the case.


I formulate this idea as "Clean code never survives first contact with users"

> Human written code just takes even longer to realize the mistakes because the pace is slower.

Yes but the ceiling is still higher, and that's the author's point. If you vibe code, without code review, code becomes a mess quickly. If humans write code by hand, then this is often the case too, but crucially, this is not unavoidable. Sure, most codebases are a terrible mess, but some are not. AIs unfortunately got trained on all of them (+ reinforcement-learned stuff) and therefore their quality standard is about as low as that of the average codebase, ie pretty damn bad.

But there are plenty examples of acceptably decent yet long-lived codebases, both in OSS and inside companies. You simply couldn't get that quality by vibe coding. (unless you review every line of code and every design decision, at which point you're about as fast as you would be writing it all by hand, assuming some seniority)


>Sad to say, but this is no different from human written code. Human written code just takes even longer to realize the mistakes because the pace is slower.

I really don't think so, poor written human code IME is rarely overly complex, where as the AI code is almost always vastly over complex. Naturally complexity can be an issue because it leads to more surface area for failures and challenges to diagnose, but where I am REALLY seeing an issue is the complexity hiding an issue. Something that should normally fail or produce an error is covered up by something multiple layers deep in the code that returns an incorrect value instead of an error when something goes off the rails.


AI written code is a function of the human created constraints around it.

That is why I believe the most senior engineers on the team with the most scars and most experience need to shift into writing those constraints instead of writing code.

In writing those constraints, they can multiply their effect across a tireless fleet of agents that generally want to copy existing patterns and can be guided to use skills.


Taste and smell still apply. You have to know the art and have comparative priors in order to judge the output.

My take: it seems like systems should become smaller, more isolated, and contract-oriented.

I have been a long time proponent of monoliths, but it seems like agents would be happier with smaller, more isolated services. The more isolated, the better. Contracts between the service components only. Then it can iterate internally as long as it satisfies the contract. If it needs to, it can version the contract and keep iterating.


Microservices are still bad. You want either a modular monolith or FaaS with most of the work in a domain library.

None of them are bad. It's like saying a bike vs a scooter vs a car is bad. Use them correctly and for what they're intended and they work fine. Use them improperly for the wrong things and suddenly people think the tools suck, when it's the humans who misused them that suck.

If we lived in a world where people or agents could define those service boundaries up front, correctly, with some reasonable foresight for change on the horizon ...then sure...I'd agree.

However, our world is not that world. Your agents may be happy in their tiny walled kingdom of toil and ineffectiveness, but you and your users will not. Poorly drawn service boundaries will drag you down more than almost any other architectural mistake.

If there is one universal amongst organizations, its that they love walls and silos. Be wary of putting them up ahead of time, cuz tearing them down once established is nearly impossible.


    > If there is one universal amongst organizations, its that they love walls and silos.
What's true for human organizations isn't necessarily true for agent-driven engineering. People and human teams struggle with contracts because there's always human negotiation involved. If the decisions are instead made by a team of agents, there's no more ego, ownership, miscommunications; just decisions based on whatever rules have been given to the orchestrator.

    > Your agents may be happy in their tiny walled kingdom...
Yes indeed; the agents will always be happier if they can iterate faster, lint faster, build faster, test faster, ship faster, with smaller context.

That would be the point of using contracts as boundaries so the agent can iterate more autonomously so long as it maintains the externally facing contract or version the contract if it needs to.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: