Hacker Newsnew | past | comments | ask | show | jobs | submit | nerdsniper's commentslogin

I got 1419 miles by putting the pin just east of Presidio, TX. The Gulf of Mexico doesn't have many Americans in it on the way out to the Florida Keys.

1,570 miles by pinning the lower Keweenaw Peninsula which null-counts a great circle route through Canada.


Having been to Presidio on a trip to Big Bend I'm not surprised

I'd love something like this for the Apple Watch but with local (iPhone) or self-hosted processing using whisper/parakeet/etc.

Fine-tuning is done against a dataset. Distillation is done against a model.

I mean they could just be routing known benchmark questions (which all of SWEBench are) to a full-performance variant.

> buying exercise equipment prescription only

Like eyeglasses for near-sighted people. Obviously prescriptions should never be needed for far-sighted people. Only for near-sighted people.


Silly complaint. Per HN guidelines as long as a workaround is available, paywalled articles are explicitly allowed.

Try this service: https://archive.ph/eJpUg


Thanks, it's working. It's more about "good manners", something posted should normally be easily accessible.

I feel like this would be much less of an issue if the datacenters were all fully solar+battery powered.

Currently it either means “better than the median human at every task including things like counting to 1000 or folding clothes” or it means “for each task, better than the top-n human at that task”.

It’s clearly not actually that “generalized” yet because it’s unable to do a number of very simple things that almost any 6-year old could do, such as count to 100 without using any tools.

It’s still a very specific type of intelligence, with some real breadth to it, but not general intelligence.


I think if you ask the median human to count to 100... more often than not you'll hear "One, Two, skip a few, Ninety-Nine, One-Hundred".

Or a seedbox on another continent.

Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.

---------------

Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]

That's a huge gap, considering that the paper was published just 2-4 weeks ago.

I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.

Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?

0: https://arxiv.org/pdf/2606.29537


You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (55.7 vs 54.8).

That is—the agent scored 100% on 20% of tasks, but on average it got 54% of the "score" awarded in the exam. One number reflects partial progress, the other one doesn't. The authors of the benchmark prefer you to look at the lower number (because they want to show their benchmark as capturing useful gaps in capabilities and with a lot of room for improvement), the authors of the models want you to look at the higher number (because they want you to think of their models as capable)


In what world is 55.7 the same number as 54.8?

What variance is acceptable to publish without a retraction?


That seems like entirely reasonable variance to me for AI models. For my purposes, that absolutely counts as a solid "replication". I'd probably accept +/- 5 percentage points even.

D'oh, they are running the benchmark themselves. Reasonable.

There is randomness in LLMs. Both papers authors probably ran the bench 1-N times. Depending on that, they might select an average, max, least, etc. They might also have discarded outliers.

Like the other person said 5% variation is probably expected


I don't know who downvoted the parent or why, but it's a fair question IMHO.

The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.

The reason is that the temperature parameter introduces random behavior.


Its slop all the way down.

I think some variance is to be expected since LLMs are typically non-deterministic, however that's a huge difference that I think warrants further explanation.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: