Hacker Newsnew | past | comments | ask | show | jobs | submit | tpowell's commentslogin

The last time i ran into that issue it suggested to make sure that a Fable AGENT took over the long-horizon task because, for some reason, agents in a session don't get blocked for security reasons. This may not always be possible, but it worked for me.


Don't fall for this. Agents silently downgrade unless you explicitly block the behavior. I built a little harness for Chatgpt, grok and Claude to do design review feedback rounds where one holds the pen and the other 2 send feedback, then rotate if no convergence. I built a thing into it to track if model swaps happen. Happens to Claude all the time. The other two, never.


Do you find the variety helps? I've migrated away from such complexity, and I simply have multiple agents of the same model run the same prompt (usually Sol 5.6 high or max), and generally this gives plenty of adversarial input. I'd be curious to know how much difference it makes to run multiple models.


I frequently find blind spots / edges where one model notices something non-trivial none of the others did. I think the one that surprises me the most often is probably grok, but I wouldn't want grok to be my daily driver. I feel I get benefits but I could also see the argument that it's just a complex token burning furnace lol.


That's a good idea but it works from the change in context no? So you lose context from your big model. It's a good suggestion though.


No idea if subagents are not blocked, but you can set the starting method for subagents to "fork", then they inherit the main agents context


Canada is (rightfully) re-thinking their F-35 commitment after the current U.S. Administration bullied them in trade negotiations: https://youtu.be/ukYxk31zXWg


Yep, I'm an American, and sadly we've proven that our country is neither honorable nor reliable. I hope Saab and others continue to provide the free world with options until my country learns that there are consequences for behaving like the slimiest used car salesmen.


It takes a bit of setup and a huge download, but every time I need a good domain I follow this old post from Derek Sivers. I have Claude de-dupe it and turn it into a searchable database (on my machine), then have it search genres and terms I'm looking for. It's a task Claude is very well-suited to, from the technical implementation to back-and-forth about selections. [link]: https://sive.rs/com


Note -- if you do this, watch out for requesting access to "all tlds". They send you two emails per TLD -- one for your pending state, and one for your approved/rejected state. I suddenly had 1k+ emails flooding into my inbox, until I found the setting on their website to disable emails.


Holy over-engineering, Batman!


I wrote this in June, and I'm honestly not sure I've felt the same magic since: I was close to maxing out my $200 plan for the week, almost all Fable use [Claude CLI]. My observations: Fable seemed to have bigger-picture thinking and completed tasks more thoroughly vs just focusing on executing the ask. It pieced together context and intent like an all-star employee would, vs one that just does what you say. Not overeager (important!), but if the above-and-beyond was warranted, it just did it. This was surprisingly delightful. Coderabbit seemed to find ~1/3 or so as many issues when reviewing, too.


This is exactly my experience as well.


I just asked Claude about defaulting to 4.6 and there are several options. I might go back to that as default and use --model claude-opus-4-7 as needed. The token inflation is very real.


I cobbled my own together one night before I came across the thoughtfully-built KeyVox and got to talking shop with its creator. Our cups runneth over. https://github.com/macmixing/keyvox/


I have an M2 Air 24GB/1TB that has been such a beast that I haven't touched my 16" Pro in months. I have four browsers running, with a ton of tabs in Brave (daily driver) and I'm sitting at 21/24GB utilization with all sorts of apps running (granted, Docker is not at the moment, but it still doesn't make it sweat). I had ~8 pro laptops in a row going back to the late 2000s, but Apple Silicon has changed how I work. A future 14" OLED that was similarly light might turn my head, but if I had to replace it today I'd just buy another M5 Air with at least this much RAM. [FYI I never installed Chrome after M1 came out. Brave has been rock-solid for over a half-decade now.]


24GB is definitely solid. 16GB is like my minimum recommended for any kind of Mac, but if you can go for more you should go for more. I think 24GB should last a good long while though.


16GB, depending on your use, can be constraining and, sometimes, you need to get creative with complex processes. My colleagues complain about developing with several containers running peripheral services. In similar situations we asked the services teams to provide mocks that answered the same APIs without needing a large memory footprint.


Can I get YoY % improvements to the geekbench scores in another column I double-dog dare you


Well I'm astounded. I talked to it for 13min, it crashed, but remembered the context when I returned a few minutes later and talked for a full 30min (it's limit).

It 99.9% felt like it performed at the level of Samantha in the movie Her.

I started asking all kinds of questions about how it worked and it mentioned a word I had to have it repeat because I hadn't heard it before: PROSODY (linguistics) — the study of elements of speech, including intonation, stress, rhythm and loudness, that occur simultaneously with individual phonetic segments: vowels and consonants. I asked about personality settings, à la TARS from Interstellar, and it said it automatically tailored responses by listening for tone and content.

It felt like the most "the future's here but not evenly distributed" interaction I've had since multi-touch on an original iPhone.


+1 for Cursor.

This guy [https://x.com/PrajwalTomar_] has been exploring workflows that involve using ChatGPT to assist in creating a Product Requirement Document (PRD), then using V0 by Vercel for mockups and bringing it all together using Cursor, maintaining continuity with markdown documents (.md) of the PRD and relevant database schemas etc inside the project to maintain continuity.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: