Hacker Newsnew | past | comments | ask | show | jobs | submit | enjeyw's commentslogin

The author hints at this but it seems like one issues is that while JEPA is good at distinguishing between unpredictable noise and predictable features, the model has no way of assigning importance to different predictable features.

So for a system where it’s very difficult to exactly reach the desired end state, the model needs to choose between (for example):

- reaching a relatively achievable scene where 95% of the features in the latent are correct, which includes stuff like visible enemies, Mario’s position on the screen etc

- reaching a far more difficult to access scene where there’s a bunch of differences in the actual level visuals, but theres a match on the latent for the tiny set of pixels in HUD that indicate you’ve hit the victory condition

We obviously know that it’s not good enough to reach an early scene that looks similar to the victory condition but isn’t. The model doesn’t.

In a sense, this is what the linear probe helps with - it allows us to re-weight the latent and say “actually, while the latent encodes many things about the world, the thing we really care about is the X position.”

I’d be curious what happened if rather than planning actions on cross entropy of a final scene, the model just tried to find the actions that maximize the predicted X value of the probe.


Hey! Author here! Yes, I completely agree with your take! I want to run another experiment around whether the JEPA can focus its distance metric on the task-relevant features. But I think this is also where world modeling starts to blend into RL or goal-conditioned learning, the world model can learn what is predictable, but an external objective still has to tell it what matters. The position probe was essentially a lightweight way of adding that task-specific weighting without retraining the representation.

Also I actually did try the experiment where we maximized the predicted X! It worked across open ground, moving Mario from x=40 to roughly x=360. But... it broke down at the first obstacle, it repeatedly chose to jump in place. The model predicted those jumps would increase X, but Mario’s real position stayed around x=342. It just couldn't get over the error in its learned jumping dynamics.

Thanks a bunch for reading :)))


I think that’s a false dichotomy.

My Lazer Genesis Helmet is a MIPs and it’s the lightest helmet Lazer made at the time.

Much more breathable than my previous helmets too.


You forgot to mention it's also $200+. Some folks buy bicycles for less than that.


The story of the navigator in the photo is also worth a read [1]. Very reminiscent of Joseph Heller’s work.

1. https://www.rbogash.com/B-52/Carls_Letter.html


What I find tricky to reason about here is that whether destroying infrastructure comes down to "whether the military advantage outweighs the impact to civilians", and as far as I can tell, there's no robust way to assess this.

Indeed, this seems to be what supporters of Trump are leaning on, as you can make the argument that _any_ bridge, or _any_ powerplant could hypothetically be used by the military, and that this conflict is sufficiently important for the livelihood of people in America/"The West" that doing anything that even slightly helps tips the odds is justifiable.


Hell, farms and water sources could be used by the military to sustain soldiers. Women of childbearing age could produce future soldiers. This line of reasoning has no floor.

There’s a reason that past generations tried to draw a line in the sand and say “we will not cross this line.” It was imperfect and often violated, but at least it served to frame actions as just or unjust. Blatant violations could catalyze domestic opposition to unjust war, as in Vietnam and Iraq. Now that the standard has been eroded into nothing, I don’t know if we can stop further escalation.


Interesting requirement. Where does that leave a lot of other wars? Russia has been attacking Ukrainian infrastructure for a while. Ukraine has been attacking Russian oil production and ports, especially recently. I seem to recall a lot of infrastructure destroyed in the US invasion of Iraq. There have been a lot of wars since WW2 and I find it had to believe that those than involved bombing were all restricted to military targets.

A lot of war is about economics and logistics.

Edit: to add, what about Iran's threats to destroy water supplies?


The problem is also that the world has made it very clear that it doesn't believe in warcrimes. Warcrimes are what's defined as illegal in the convention of Geneve.

The purpose, the idea behind warcrimes is that when warcrimes occur, the world would unite, in the security council, a mandate would be voted in, and the whole world would intervene, preventing warcrimes from occuring, or at least from repeating.

Well, when it comes to Iranian and US warcrimes the UNSC, specifically France, Russia and China have declared there will be no consequence to any warcrime by either side. In France's case it's not that they don't think warcrimes are terrible crimes, it's that they don't want to help anyone.

In Russia and China's case it's that they think this war destabilizes the west and that matters more to them than terrible crimes. Oh and the whole communist stick of "it's not warcrimes, it's internal matters", you know, when they do it to their population. So they have declared they will actively fight to prevent anything being done about warcrimes.

Under those circumstances, of course, warcrimes effectively don't exist, and that's that. Or to put it another way: the world is perfectly happy for you to be discussing the finer points of international law and why this and that is or isn't "a crime".

But the world is totally unwilling to do anything about warcrimes. I mean, let's be realistic. The world is unwilling to do anything about Iranian warcrimes, and perfectly certain the US won't commit any (the US will make mistakes, of course, but not actually commit real warcrimes). Whatever the outcome of your discussion on what is and isn't a heinous crime ... there will be no consequences.


Hypothetically civilians can be used by the military and provide some military advantage as future soldiers or weapons manufacturers or even army rations providers. So let's bomb them too, right?


One thing to consider is that Trump is publicly stating that the US are destroying the infrastructure as a punishment for non-compliance. That basically makes it clear that the motive is not based on military considerations.


It's basically what russia is trying to do for years in Ukraine. Beating populatuon into submission. Which is even dumber in case of Iran since it's not a democratic country where population has much of a say.


Yeah I did wonder myself if that tweet was an admission of guilt.

If I were a lawyer responsible for defending Trump in the Hague, I'd argue that the tweet was actually an abbreviated way of saying "If Iran does not comply, we will destroy all military assets, including but not limited to their ICBMs, Bridges, and Power Stations, such that we have total military dominance."

Now very obviously (to me at least) this was not the intent of the message, but I don't know whether you could prove that in a hypothetical war crimes trial.


> “…tips the odds is justifiable.”

The slippery slope.


exactly, they can argue forever that their point of view was justified.


Ok I’ll bite…

37 miles?!? Why??


Land-mobile radio stuff. Analog, voice communication.

Our sales guy had sold a remote node for a voter system to improve receive coverage for a central dispatch system. (Signal-to-noise voters are pretty neat: They can continuously compare two or more related audio signals and [ideally!] pick the one that is best for use while discarding the others.)

That node wasn't all that far away as the crow flies, but it was a very long way out in telephone cabling miles. It spread across two different telco LATAs.

So we rented this very long series of bits of wire held together by scotchloks and punch blocks and whatever else in telephone world to use, and we used it. It was not a conditioned circuit: Just wire.

The specific endpoints of that wire were kind of neat, too: There was some basic EQ that could be used to help compensate, and (IIRC) some impedance adjustment to dial in the circuit itself.

And there was a continuous pilot tone used to set gain: Apparently, when wire gets really long like that, atmospheric conditions can dynamically change its attenuation.

Putting a pilot tone near the middle of the voice range (to be notched it out later) and using its level to set gain helps to improve consistency.

That wireline stuff all worked pretty well.

(The remote node was ultimately a bust. The sales guy also tried to cram too much shit into one feedline and antenna, and the gear to combine and separate all of those signals ate too much energy to make any of it an improvement over doing nothing at all.

Which is... well, that's exactly what the engineering told him would happen, but he did it anyway.

No part of this was inexpensive.)


One of the big problems with Attention Mechanisms is that the Query needs to look over every single key, which for long contexts becomes very expensive.

A little side project I've been working on is to train a model that sits on top of the LLM, looks at each key and determines whether it's needed after a certain lifespan, and evicts it if possible (after the lifespan is expired). Still working on it, but my first pass test has a reduction of 90% of the keys!

https://github.com/enjeyw/smartkv


Is this not similar to DeepSeek lighting indexer


It absolutely has been!

In general prediction markets can’t be “correct” or “incorrect” - for instance if a prediction market says there’s a 60% chance of an event occurring, and it doesn’t occur, was the market right or wrong? Well it’s hard to say - certainly the market said the event was more likely to occur than not, but only just, and who knows? Maybe the event _only just_ occurred, and very nearly didn’t!

So generally we say a prediction market is “correct” if it is “well calibrated”, which is to say that if we took all the events that the market said had a 60% chance of occurring, then approximately 60% percent of these events occurred (with the same holding true for all other percentages).

On this note, an interesting phenomenon that used to occur was “favorite-longshot bias”, where markets would consistently overestimate the likelihood of longshot events occurring - so events that the market predicted would occur 10% of the time would only occur 5% of the time. What’s fascinating is that once people realized that this bias exited, they began to exploit it by making bets against longshots, which had the effect of moving the market and removing the biases, making the markets well calibrated. It’s a pretty neat example of the efficient market hypothesis in action!


Some of the longshot biases still exists and can't be removed due to technical constraints on the platforms. A lot of times there is a minimum contract price, which effectively means the probability of unlikely events cannot be modeled as lower than 1% or 0.1% or whatever. But there are contracts for events much less likely than that.


There are also issues with the time value of money for long-shot events. Someone has to be willing to buy a share of "No", and if that works out to a return lower than the risk-free rate (eg. buying t-bills) there will be no incentive to take the "No" position. That makes anything roughly under 3-4% per year pretty unreliable.


Polymarket and Kalshi both pay interest on long term bets around the same as the risk free rate.


> for instance if a prediction market says there’s a 60% chance of an event occurring, and it doesn’t occur, was the market right or wrong? Well it’s hard to say - certainly the market said the event was more likely to occur than not, but only just, and who knows? Maybe the event _only just_ occurred, and very nearly didn’t!

For most events like this, you'd want to see the market spike to 0% or 100% as the deadline approached. And in particular for an event that happens, you want to see the spike to 100% before it happens. Remaining at 60% until after the fact is wrong because the occurrence of the event becomes more certain as it gets closer.

Being "well-calibrated" as you describe is a very bad quality metric in the sense that two sets of predictions can achieve the same calibration profile while differing markedly in quality. The farther the predictions are from 50%, the better they are, but your calibration metric doesn't take this into account.


Charlie Kirk has a 3% chance of winning a Nobel Peace Prize right now according to Polymarket. He's climbed from 1% since Maduro was arrested.

It seems unlikely since Nobels aren't awarded posthumously.


The issue there is time. The Nobel prizes will be announced in around 9 months. Buying a share of "No" would currently cost 98.2 cents, working out to a (basically) risk-free return of around 2.4%. Alternatively someone who wants a very low-risk investment product could just buy 1-year t-bills with a return of... ~3.5%. And that doesn't require messing around with buying crypto and the inherent risk of trusting Polymarket with your money.

Anything under 3%/year of time until decision is going to have pretty limited predictive value within that range. Anything starting above that range will end up hitting that floor rather than going to zero because of the difficulty of finding a counterparty.


Are t-bills accessible for anyone, especially loaded with crypto?

AFAIK there are coins that pay better (some give you exposure to t-bills).


Polymarket pays interest on those bets about the same as the risk free rate.


It doesn't, or where do you think those 3% are coming from?


While Polymarket does offer holding rewards interest, it looks like it doesn't for this particular market.

That doesn't mean there aren't other explanations. It could mean that No holders expect to incur an opportunity cost greater than the risk free rate. Combine that with how there's low liquidity (there's less than $300 on the book buying Yes, and at 2 cents or less), and so we could just be seeing the effect of random fish temporarily distorting the price. It could also mean that the risk of a smart contract failing is making it not worth the hassle for a market maker to come in at such a slim margin and low volume.


They're offering interest on roughly a dozen hand-picked markets, according to their documentation. (I wasn't aware of that, so I stand corrected on the general assertion that they never do.)

> That doesn't mean there aren't other explanations.

Why do you need other explanations, when the observed probability can be precisely and fully explained by opportunity cost?


I don't have to "need" other explanations in order for them to exist. The current price does happen to accurately reflect what the risk free rate would imply. But look at the graph history: it hovered around 1% for a large chunk of December.


How much volume on this bet? Let's ignore black swan events and say it's a guaranteed 3% return. On how much? $1? $10? $1m?

I'd weigh the accuracy by how much money is at stake...

Even then, a "perfect" prediction market need not be accurate, if people use it for hedging. If some low probability event is really bad for me, I may pay over odds (pushing the implied probability up) to get paid if it happens. The equilibrium probability may be efficient, reasonable and biased.


Nobel Peace Prizes are Peace Prizes as much as they are Nobel Prizes.

I'm not sure the same(any) rules apply.


They would be a better prize if the word ‘Peace’ was removed from the title.

It would make the likes of Kissinger getting it easier to understand.


Well, Nobel peace prizes aren't usually awarded to people calling for invasions of their home country either, or cheering for the extrajudicial double-tap killing of smugglers/random fishermen.

Who's to say a dead person can't have done the most to "promote peace conferences" as mentioned in Nobel's will? These days, I'd say dead people make a larger net contribution to peace than most politicians.


Normally they aren't, but maybe the US will take over Sweden and the Nobel Foundation and make it happen.


...only to find out they invaded the wrong country! (Nobel peace prizes are awarded in Oslo)


To be fair, it would bw totally on brand. They would not admit mistake tho declare it a success and award own nobel price.


Ooops. I thought something was off when I looked it up (headquarters of Nobel Foundation are in Sweden).


To be fair you’re not really providing a hard stance against the estimate. You say it is unlikely, and indeed the prediction is a 3% chance. That’s unlikely.


No, markets are evaluated on accuracy, not calibration


Well markets are evaluated on a number of different metrics depending on what you’re trying to determine.

If you want to go be pedantic about it and select one metric, markets are evaluated on their Brier Score or some other Proper Scoring Rule, not accuracy.

However, I prefer calibration as a high level way to explain prediction market performance to people, as it’s more intuitive.


Yeah it's a good way to introduce the idea. But I don't think someone would really grasp it until they understand why both calibration and "discrimination" are necessary in determining if a prediction market is accurate.


Proper scoring rules measure accuracy


I suspect that you are arguing semantics, where parent and grandparent focus on the nuance of what is ACTUALLY being measured. I am saying it like this, because while I never used prediction markets, I briefly looked into them to see if I could use them well. The question of accuracy came up, which is why I happen to align with posters above.

With that in mind, what do mean exactly.


Noob question from me: what’s the difference between accuracy and calibration? A well calibrated market would be more accurate and vice versa, not?

Edit: just found the answer myself: “accuracy measures the percentage of correct predictions out of total predictions, while calibration assesses whether a prediction market's assigned probabilities align with the actual observed frequency of those outcomes”


Suppose there are 1000 events and 500 will have outcome A and 500 will have outcome B. If you predict a 50% chance of A for every event you'll be perfectly calibrated. On the other hand, if you predict a 90% chance of a certain outcome and you're right for 800 events, you're not perfectly calibrated but you have a lower Brier score (lower is better).


A forecaster can be calibrated but almost only assign probabilities in the 40--60 % range. This is not as ueful as one assiging calibrated probabilities in the full range.

We try to measure the increased usefulness of the latter with proper scoring rules.


I used to share a somewhat similar sentiment.

I know one anecdote is not data, but his investment in BYD all the way back in 2008 does counter that viewpoint somewhat - his investment success in the BYD case isn’t from other investors following him in, it’s from him identifying BYD as a successful company far before any other major investors did.


Minor nit - it was Charlie Munger who identified and argued for BYD.


Oh I didn’t know that! Thanks for the clarification.


Overly specific LLM research into KV cache eviction.

The vast majority of tokens in a sequence will be irrelevant to an attention mechanism outside of a very small window. Right now however we tend to either keep all cache values forever, or dump them all once they hit a certain age.

My theory is that you can train model to look at the key vectors and from that information alone work out how long to keep a the token in the cache for. Results so far look promising and it’s easy to add after the fact without retraining the core model itself.


I made a tool for this! It's an essay writing platform that tracks the edits and keystrokes rather than the final output, so its AI detection accuracy is _much_ higher than other tools: https://collie.ink/


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: