Hacker Newsnew | past | comments | ask | show | jobs | submit | foota's commentslogin

I feel like multimodal models that can read images should work differently than they do. My understanding is that multimodal models basically first generate an image embedding and then the model is trained to interpret that embedding, but in the same way that text is lossy, it seems like the embedding would be as well. Why don't multimodal models learn to interpret images themselves without an embedding? Or e.g., by passing some "prompt" to the embedding model?

What does interpreting images mean in practice if you exclude the possibility of feature extraction or any other sort of implicit embedding?

I'm not an ML expert, but I was thinking of a sort of "guided" embedding. E.g., give the image model some prompt for what it's trying to do? I don't understand why multimodal models generate an embedding that doesn't understand what the model is trying to "figure out".

I think this is similar to how Gemma 4 12B is implemented, but even then I don't think the single layer image embedding is "aware" of the context.


Jokes aside, I think the idea is that the law is simple to code, the proof that it holds is where the agent is responsible. This probably becomes less true though as you try to express more complicated laws.

Heh, it's like we all need to collectively read I, Robot yet again, and the myriad of SF books on the subjects. Black and white quickly dithers to grey.

Indeed

Robots are logical, but not rational.


The point of those books is that robots can be perfectly rational, and for a useful robot we would expect them to be.

What they aren't is moral, because of the orthogonality principle: you can't use facts and logic to discover correct moral beliefs. Morality is about values, goals, and the definition of "good". They must be provided to the robot by its creator, and those are things that are very hard to precisely describe in a way that is fully consistent with the speaker's intent in all possible scenarios, and agreeable by all other people.


LLM driven agents aren’t even that.

A generation is a time that people were born between and useful because it embodies much of the cultural context that people share.

No, it doesn't. The baby boomers represented an actual demographic change, that's why they're tracked by the Census Bureau. Everything since then is marketing and astrology.

you are just nitpicking.

People get used to understand what generation 'name' they got from society, thinktanks and co and its not that hard to know what the younger generations are.

https://en.wikipedia.org/wiki/Generation#List_of_social_gene...

Astrology is just garbage.


Named generations are a marketing tool, it’s a concept as garbage as astrology

Its a name for a generation a decade wide. I really don't get your hate for it?!

It's the same, you're just substituting years for months.

The difference is that named generations don't claim that any differences are due to supernatural influence. Instead they are arguing that the differing circumstances in which people grow up and live influences their experiences and how they see the world. That is a MUCH more defensible position.

It would be defensible if it worked that way. Instead they work backwards: define the groups for marketing purposes, and then try to identify the similarities.

That's irrelevant to what I wrote. Whether it was made up by marketers or anyone else doesn't change the fact that unlike astrology it doesn't make an appeal to magic.

The name of the generation gives you birth years. It also tells you "hey the generation 2010-2020 are getting hit by AI a lot more than the previous ones.

This not astrology at all. Astrology is 'magic thinking'.


BS. While I agree the cutoffs are arbitrary, marking a cohort by, say, "people who graduated highschool before the advent of smart phones" or "people who went through childhood in the early years of social media" are valid groupings.

As a Gen Xer my childhood was soooo different from what kids today experience. And literally every single Gen Xer I know is glad and feels relief (sometimes mixed in with remorse) that our childhood was before the consumer Internet took off.


If you want these cohorts to have statistical significance you have to add more details, such as the geographic location, language, etc. But you can create cohorts with pretty much anything, in itself that doesn’t validate the concept of generations as pitched by marketing agencies

Luckily for us, the cohorts don't need to have statistical significance because people aren't talking about statistics. They are talking about the broad shared experiences of most people born in a given range of years.

Funny enough I was thinking about something very similar to this based on the Jev model posted yesterday.

I’ve played with it already. I don’t think this is the use case. I think Jev’s use case is fast, cheap and somewhat easy classification. It’s not trainable in the way you would want here. Even though it’s fast it wont be faster than pgs query optimizer.

At least as I understand things.

How did you plan to use Jev for query optimization?


I am still struggling to understand a use-case for Jev. Isn't what was explained in this article a classification problem? I.e. find and aggregate data?

The number of options has to be small and bounded. The query planning is more of a search/optimization problem than a classification problem since the number of options increases wildly based on query size.

It’s classifying faster and cheaper. A lot of immediate ideas are better solved by pre-classifying + embedding, but their doom example or the wikipedia runs are one where you can’t preclassify.

Just curious, why? Is this true even if you did something like a per-CPU histogram that uses atomic ops to increment?

If you have a per-cpu metric there would not be a reason to use atomic instructions to mutate it.

In general your unpinned userspace threads will hit the same CPU 99.99% of the time, but not 100%.

Sure. You get the pointer, you lock the mutex, 99.99% of the time that is uncontended, then you set all the metrics and release it.

Taking the mutex uses (uncontended) atomic ops.

You got an article off by one error, I think you meant to post on https://news.ycombinator.com/item?id=49558685 :)

Ah this is interesting. Essentially the idea is that the compute can try and move into positions that it can evaluate but humans might have trouble evaluating because of the board state's complexity?


That's the basic idea, yeah.

When playing white in handicap games, you want to make your opponent uncomfortable.

Play moves where the simple/safe/obvious move is just a little bit bad. Force them to choose between complex fights or a slow death of 100 slightly suboptimal moves.

It feels really wrong to defend like 10 times in a row, so if you make them do that they'll lash out at the wrong time and you can take advantage.

You also want to look for moves where...even if their best response means it's even or a little bit worse for you, there's ~reasonable responses where you win out or it goes complex.

A lot of the time it's not even crazy complex fights, it's more just situations where the judgement of what is more points is difficult.

(Note: most of this stops applying as strongly if it's a teaching game, which most handicap games are, there you have other considerations besides winning)


By the way that is exactly what humans do when playing with white in high stone handicap games. They place their stones all around the board, start little fights everywhere and wait for the weaker player to misread or misevaluate something. Suddenly two or three fights merge in a one sided larger one and part of the handicap is gone.


I think they meant more like "people band together to pay Christophe Nolan to create some new film" ala crowdfunding. Is that a sustainable model for the world economy? Unclear.


Well if we got enough people together it would be pretty cheap, maybe 15 or 20 bucks or so with inflation. And we could fill the movie upfront, and create some previews and launch an advertising campaign -- that'd certainly help drum up interest and get more people to crowdfund. Best even to just collect their money as right before they watch it so they know they aren't getting scammed and will in fact get to watch the movie they are crowdfunding. We can also crowdfund some popcorn machines and communal viewing rooms, after all people will want to have a snack and watch their crowdfunded movie with friends and family.


That was a lot of snark :-) I know what you're saying. It's not clear to what degree the film industry would be viable in its current form though without IP rights. If everyone can watch the latest marvel film at home for free, are they going to be willing to pay to see it in theaters?

There would likely still be some demand for the experience, but the tickets would need to be more expensive to recoup the costs of making the film. At that point: it does seem like people pooling together their funds to fund the movie itself isn't insane. The point is that the "investors" wouldn't be investing in the hopes of a financial return, they'd be paying for the creation of the film.


> Ban offline training/pretraining. Models must train from scratch after submission Previously this was considered impossible so rule. My model shows this is possible Guarantees no synthetic data can be used It makes the comparison fair across differet models. Otherwise some models like LLMs can benchmaxx ARC by using ungodly amounts of offline training. (Since the benchmark has been around a long time, many ARC-like datasets have been created)

I'm not an ML researcher, so YMMV, but... how could a model learn to answer these ARC-AGI questions without training beforehand?


train from scratch only during the 12 hours allowed on Kaggle

Other competitions have implemented things like this before. Eg: OpenAI's Parameter Golf and Keller Jordan's Modded NanoGPT Speedrun


how would you as a human know the answer to the arc-agi questions?


I spent years learning logic and doing puzzles. I don't think a baby or even average kid could solve these.


Predicate logic is trivially realized by linear transformations (I.e., matrices), and these matrices are easily discovered via gradient descent with appropriate reward functions.

> I don't think a baby or even average kid could solve these

The reward functions of a typical baby or kid is not 'get a huge dopamine boost when you solve a logic puzzle' (or whatever neurotransmitter, I don't know).


> I agree that its rare to see to face problem sets in real life where every problem is given at once. Even if it is (like an exam), humans can usually only attempt one at a time

Just one small snippet that I thought was interesting. I would always read through ~the entire exam before starting. Both so that I could find the problems most approachable to me, but also because sometimes it helps me figure out the rest of the questions :-)


same! I'd often learn during the exam by solving an easier problem and then that would let me tackle a hard problem


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: