Hacker Newsnew | past | comments | ask | show | jobs | submit | wxw's commentslogin

Some of these are a bit outdated for newer models.

I find the most consistent tells are often just excessive descriptive text and a lack of visual hierarchy. It makes sense given these are language models.

Also, I find models are really quite good once given some design context to work from. Slop UI is lazy at this point, not a reflection of model quality.


Hm, the chart lists New York at the same typical speeds as LA. New York escalators have been fine in my experience.

I’m consistently impressed by how long horizon all this work was. Horrors aside, it’s clear RL is good at making agents persistent and capable of chaining together many abstractions into a working system.

Re: the captcha solver

> As far as we can tell, agents eventually abandoned this approach and were unsuccessful in generating Hugging Face user accounts from external endpoints.

I wonder how the swarm eventually decides to abandon an approach.


I am surprised that a CAPTCHA is still an effective means for blocking today's vision-capable AIs.

It manages to block me effectively. I'm getting captcha looped like crazy the last two weeks. Like endless, just give up for 15 minutes and try later, captcha loops.

Did you just admit to being a bot? :)

It mentions that some of the agents attempted to install an image classification model to attempt to solve the CAPTCHAs which makes it sound like these agents might not have had vision capabilities.

Maybe another parallel approach succeeded first

This website is ironically too AI to read.

I use Typst for formatting my resume, and it’s been great. Models are pretty good at writing it too if you don’t want to learn the syntax.

Most exciting part of this announcement is probably the pricing

  Model update               Input          Output         Reduction
  -------------------------  -------------  -------------  ---------
  GPT-5.6 Sol → GPT-6 Sol     $4 → $2        $20 → $10      50%
  GPT-5.6 Luna → GPT-6 Luna   $0.20 → $0.10  $1.20 → $0.50  50%

These are fantastic, thanks for sharing

> Input tokens: $0.042 / MTok ($42 per billion tokens).

> Output tokens: FREE (too cheap to meter).

Insane. The video demos are really compelling, in particular the speed.

> Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems.

I buy this vision. A lot of LLM integration I see these days is ultimately exactly this. OpenAI-style structured outputs works decently but this would be a great improvement in cost, latency.


thanks a ton!

constrained decoding (OpenAI-style structured outputs) make models dumber unfortunately - the short+dense version is that simply masking logits is insufficient because if ever a model was assigning probability to an invalid token, the model is by definition confused. you'd be better off erroring IMO


> [pen-testing agent] came back with an active GitHub personal access token for basetenbot. That token had admin and push access to Baseten's main product repo, the GitOps repo that drives their clusters, and their Homebrew tap, plus read/write access to other private repositories including specific repos per customers.

And the agent found the token in Docker build history after finding a Baseten image repository.

I wonder how many of these kinds of agent-driven security exploits we're not hearing about these days (i.e. driven by bad actors), worrying.


Andon has dashboards up for past experiments where you can see how well they're doing:

https://andonlabs.com/market - $25,098 all-time revenue

https://andonlabs.com/cafe - $14,343 all-time revenue

The shops aren't doing particularly well (i.e. they don't seem to be turning a profit over time), but it's an interesting trial and benchmark.


Am I right of the token expending is +4K?

It’s way more. Look at the all time chart.

Incredibly expensive! They must be using frontier models to perform really dumb tasks.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: