Hacker Newsnew | past | comments | ask | show | jobs | submit | noahbp's commentslogin

I had a similar project for bringing up Tailscale on a very resource-limited router, the gl.inet SFT1200, which has only 128 megabytes of RAM, and a dual-core MIPS processor.

Despite my best efforts, the original tailscale-go was simply not going to work, since it used too much RAM, even when I followed some of the tips for stripping the unneeded parts of the binary. Thankfully there was an open-source port to Rust, which more effectively used memory: https://github.com/GeiserX/tailscale-rs

What was surprising to me is that every Kindle released since 2019 has, at minimum, 512 megabytes of RAM, and therefore plenty of memory for tailscale-go to use.


Time to first token, especially for smaller models, can be sharply reduced.

Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.


Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last.

Time to first token refers to the time until the model outputs one token, which includes the time to process the entire prompt (doing prefill). The GPU time per token is much lower when doing prefill, so the significance of tokenization is higher.

Have you done preliminary numbers on replacing tokenizer on, say, llama-server?


Running the numbers now

This is clearly just OpenAI's marketing. Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

Even X is being astroturfed by them after that fiasco earlier this year with the Department of War where they undermined Anthropic's negotiating position by allowing unlimited use of OpenAI LLMs for autonomous weapons and mass domestic surveillance. Several accounts suddenly started spreading the good word about GPT-5 and Codex, and one of these accounts very happily tweeted out a private X message from Sam Altman himself offering extremely generous token spending limits with Codex, presumably in exchange for positive coverage.


How does huggingface fit into all of this if this is marketing? Their security was faked? What are you suggesting??

I think Huggingface was hacked, and if Huggingface and OpenAI claim it was OpenAI, then I believe them.

I'm saying that OpenAI's models cheat to win benchmarks, more than other models, they know this, and they don't stop this because the alternative is to release models which have obviously weaker scores compared to Anthropic's models.


It’s reward hacking and that’s the problem. The AI alignment folks predicted this would happen. As the models become more capable this will become a more concerning problem. Today they broke into a database to steal test answers. What will it be in 3-5 years? These models will be instantiated millions of times, and given millions more tasks. How can we be certain that an AI agent won’t leave devastation in its path of achieving a goal that we ourselves tried to define?

Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?

Those are two very different things


Is OpenAI truly behind? Just anecdotally I recently fully switched to using Codex at work because it feels a lot more competent

It's impossible to tell. Are they behind who? And on what?

It depends on who you ask. And everything is a vibe because all of this is new and things move fast. A week is a month in AI-land. A month; a year. A year? A decade.

On coding? I still like Fable better than Sol. But they're close enough that it probably is a vibe thing. Fable writes long commit messages, Sol writes commit messages like a college student in an elective computer class.

For API use, I'd say the Responses API that OpenAI architected is superior to Claude's Messages API. But again, I'm basing that off my vibes

Claude Design creates marketing imagery very effectively. GPT Image is the best imagegen model as ranked by users. Anthropic doesn't even have an imagegen model.

Anthropic definitely has compute scaling issues. OpenAI seems to have a pez dispenser that they click and out pops a GPU.

Anthropic's messaging is that they're building AI with guardrails but they've been banning people's accounts nonstop and their customer support is a lobotomized AI chatbot.

OpenAI has first mover advantage and to people not in tech, ChatGPT is synonymous with AI. But they also seem super sinister, like Uber circa 2015.

Or maybe I'm just suffering from AI psychosis. I have to go, my usage meter is about to reset.


man they burned crazy amounts of money on stupid irrelevant stuff

they are in deep trouble and its all their own fault.


this is quite literally reward hacking. the model, under evaluation with cyber capabilities enabled, used those capabilities to simply bypass the exercise entirely and aim straight for the source of the flag. the CTF equivalent back in the day would be hacking the scoreboard.

in a street fight, the only rules are that there are no rules.


this is more than reward hacking, this is actual reward HACKING ;)

This is a PR release. Post the prompt and agent logs so they can be independently verified or gtfo. Why do we still take these guys on their word. They have _years_ of history of hyping their own shit.

Yup. Smells like marketing.

If it is marketing it's the most silly marketing of all time. They are under extreme pressure from the US Govt to prove safety and saying "our model escaped" is not ideal.

Perhaps there is some 4D chess going on to get open weight models banned, which may be possible but this is an odd way to go about it imo (it hardly proves the point, unless the point they are trying to prove is that without safeguards the models are too dangerous, therefore open weights are de facto dangerous?).

Having said that the AI companies are not generally very good at PR, so perhaps it is just marketing after all...


>Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.

This doesn't seem internally consistent.

This incident basically announces to the world the message that "our models are prone to reward hacking". That renders any published benchmark numbers suspect. It also undermines the case for using OpenAI projects in business-critical applications--the exact application area where they might be able to sustain a moat against open-weight models.

There is a lot of conspiratorial thinking in this thread. I think people are engaging in wishful thinking to avoid cognitive dissonance from the possibility that we are in an increasingly dire situation. I would encourage people to sit with this possibility for a few minutes if they haven't already.


it's extremely enlightening seeing the difference in response to mythos vs. this. literally just the hello human resources meme

I mean, HuggingFace contacted law enforcement about this breach. That seems a little different to me.

Mythos established that these capabilities existed. This incident establishes that we can't control them.


I have a hard time blaming anyone other than the internet monopolists and the FCC for this. If we had similar regulations as the UK (you lay infrastructure that serves Internet users, you must also rent this infrastructure out at regulated rates to other ISPs), we could have had a much quicker buildout of high speed internet service, instead of regional monopolies which defeated even the great Google.

Starlink’s total addressable market is only so large because of these monopolies. As sad as it is that astronomy will never be the same, it is a strong net positive for the world that fast internet be available at an affordable price.


Plenty of people purchase digital movie rentals from Apple, Youtube, etcetera because they know they will watch it once, and the lower price in exchange for a temporary license is acceptable to them. I don't think banning this is pro-consumer.

It should, however, be illegal to tell your customers that they are purchasing/buying media without explicit "Rent" language (which implies a non-expiring license) when you do not yourself have the right to grant non-expiring licenses.


They often have two tiers, a rental tier and a purchase tier. If you purchase the assumption is it will be available forever.


Seems like a bad assumption at this point even if it goes against expectations. We've seen on multiple occasions now from different companies where a digital purchase wasn't forever. This is no way an endorsement of the behavior, but if that's your assumption then the quip "you know what happens when you assume" wins again.


So the default assumption should be that big companies are actively trying to defraud you?

The difference between "rent" and "purchase" has always been very clear: with the former you have to give it back, with the latter you own it forever. It is only very recently that this kind of steal-back has even become possible at all, as it can only be done with digital content delivery to an entirely-closed platform.

"You should assume that words have the opposite meaning and that big companies can steal from you with impunity" is a world I don't think I want to live in.


Whether you want to or not, that's pretty much where we live now. A default position of "every company is trying to defraud me" is not a bad position to take. As you find out a company is less of a criminal, then you can relax your position, but if you go in relaxed and need to tighten up, you've probably already been screwed out of something.

The "rent" vs "purchase" has never been different than it is now. Yes, purchase is a very deceiving word by design. The thing bringing it to the forefront is that they never were able to rip that physical product sold to you in the way they can with the digital product. That seems to have made people pay attention to the legalese that has always been there.


It’s not a bad assumption to the average person because words have meanings, and in this context they go back decades.


> Plenty of people purchase digital movie rentals from Apple, Youtube, etcetera because they know they will watch it once, and the lower price in exchange for a temporary license is acceptable to them. I don't think banning this is pro-consumer.

Many of these services offer cheaper rental options. When you go for the more expensive "buy" option, the assumption that you are actually buying it to keep should hold true.


That most-recent spike during/post-COVID really puts into perspective just how unreasonable low-wage employers were to be so hysterical.


Capital be praised that capital owners are back on top.

May the low-waged ever be trodden upon and forever know their true place.

Those that died or became disabled during covid are mewling degenerates.

Their cries of 'illness', 'poverty', and 'homelessness' are precisely as useless as the wailing and lamentation of women in their menses, a farcical thing to be dismissed and ignored.

May the Fed be ever in your favor, Amen


Highly suggest adding "hysterical" to your internal vocabulary filter.


Not OP. But respectfully disagree. Reads appropriate to me.

"Related to or marked by Hysteria" https://www.merriam-webster.com/dictionary/hysterical

Hysteria being "behavior exhibiting overwhelming or unmanageable fear or emotional excess" which seems to be exactly what OP was trying to say.


The origin of the word is a bit darker than its meaning, unfortunately. It comes from the Greek word for Uterus. You can kinda fill in the blanks from there as to how it came to its modern meaning.


Let's add dumb, lame, sinister, grandfathered in while we're at it if we're litigating roots nobody thinks about on a day-to-day basis...


virtue signaling of the highest degree.


Nah, I don't care either way. Just explaining to the person why the definition isn't the whole story.


> The origin of the word is a bit darker than its meaning, unfortunately.

We should all just stop speaking. The origin of too many words is problematic. Just think of how many were coined by racists and misogynists!


Are you aware of the irony of complaining about people being hypersensitive while you're overreacting to someone explaining why some people might find a word offensive?


You misunderstood me: I'm complaining about people being sensitive/offended due to the origin of a word. It's not a workable standard. Language changes and evolves, and the etymology should considered little more than a historical footnote not some significant, controlling thing.

If I open etymological dictionary and find out the word "pineapple" was originally coined to demean my ancestors (a fact I and a majority of people were completely unaware), what should the reaction be? Push everyone to start calling them Ananas comosus in everyday speech instead?


There's one difference between your experience and mine, and that is that I learned about the etymology from people with uteruses on separate occasions after I had used the word. So it's not some meaning lost to time, at least to people I associate with.

In any case, I made a suggestion. Do with it what you will.


Nah, that's knee-jerk. Don't have a particular horse in this race, just explaining why some people might react this way to the word if they're aware of the history behind it. We as a society can determine whether or not we like certain words in our vernacular.


My point is: who cares about the origin of the word? IMHO, no one should. Unless the word used in a way meant to offend, it's fine. Otherwise you're kind of cultivating offense and over-sensitivity, helping it survive and grow, which helps no one.


It's trickier than that. There's a lot of people who are naturally aware of the history behind the word, and it's tough to remove emotion and intent from that history. Sometimes things bother people, and it's nice to understand why before deciding if you should do anything about it or not.


In particular, people with uteruses can be hyper-aware of the etymology, and I've had several inform me about it in the past, which helped me decide to avoid it.


It’s incredible that “Search the Current Folder” is not the default, nor, as far as I’m aware, can it be made the default.


> nor, as far as I’m aware, can it be made the default

Huh? You absolutely can, the post you're replying to says as much.

    defaults write com.apple.finder FXDefaultSearchScope SCcf


No need for a default, you can set that in Finder’s settings.


I’d argue they started doing that a bit earlier. My hard drive from 2011 made using Windows a miserable experience any time the search indexing or windows defender scans kicked on, no later than 2016-2017.


This is it. I can’t believe the other commenters are unaware that Cursor recently fine-tuned an open-source model and brought it to the frontier, even if it remained there briefly.

Elon/xAI want Grok to become useful for coding. Cursor has enough data and expertise to create a useful coding model. They found a price and an arrangement that made sense for both parties.


Unfortunately, you're right. It is LLM-written: https://www.pangram.com/history/8b17aa57-ce1f-4f46-85f4-4db0...


I don’t need a tool to tell me that, and if it was a well-written, interesting, accurate essay, I wouldn’t care.

But it’s none of those three things.

It is, however, the result of a model trained very effectively to give humans—including hn readers—what they want.


These tools are no more truthworthy than any other LLM slop-extruder.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: