•
>Meanwhile, human review and comprehension are starting to fall behind. For example, people are still involved in the "archeology" of the OpenAI-HF incident from many months ago. Mathematicians may be poring over the 722 manuscripts on frontier mathematics for a while. Amid all the discussion of sigmoid curves, and where the "LLM wall" will materialise, I think few people would have predicted that the real wall in LLMs would end up being humans' capacity to verify the output. What I fear is that people simply eschew human review altogether, considering we're talking about the industry that came up with the "move fast and break things" credo. Human review of LLM-produced code where I work is already a farce, and we're not special enough to be one of Karpathy's 5,000. I do my best to manually review anything that's my responsibility, but I'm literally one of very few people left working on my team, so in practice what happens is I submit PRs that are at best glossed over by completely unrelated teams for security, malware/prompt injection, and other serious concerns. Quality insofar as vetting others' code has completely gone out the window and it shows in the number of bug reports that come back, often themselves written in Claudease. Worse yet all the incentives point to this being the most economically viable thing individual companies can do. I think it goes without saying some type of regulation here is urgently needed, and that an unexpected cause of an AI bubble pop may end up being that humans simply aren't able to keep up with the pace of the output - leading either to precautionary plateauing of capability, or major liability risks related to a decline in quality.
••••••••••••
There seems to be two common fallacies in SV technocrats' AI discourse: (1) A far-reaching tendency to overextrapolate from the low-hanging fruit of the last few years of pretraining progress. GPT-2 to GPT-3 may have been a quantum leap, but GPT-3 to GPT-4 was not, and GPT-4 to 5 even less so. The party has been kept going by RL and agents, but still, there is indeed a point of diminishing returns, not just relative to available compute but to how much training is possible when the entire intellectual output of humanity, plus a raft of synthetic data, has already been inhaled by the training process. If one is to internalise the things that are said here on HN with regularity about model progress, and sentences ending with "yet" or "for now", then it would be easy to conclude that my 10 year-old son, who gained 3 inches of height last year, will be tallest structure on the planet by age 17. (2) Inability to distinguish between technological, computational, and energetic limits of LLM capabilities vs. ontological / conceptual ones. There are some things LLMs cannot do, or at least do well, at any size, at infinite size and with infinite compute, simply due to the very nature of what LLMs are to begin with. This latter topic receives almost no attention, except maybe from Gary Marcus and Yann LeCun. In that respect, this article is a breath of fresh air, insofar as it highlights that LLMs aren't "AI" at all, as we have traditionally understood the concept. They really _are_ stochastic parrots. The relevant questions are about how much that matters for some domain or set of applications, not whether they are an emerging alien intelligence with civilisation-threatening capabilities.
•••••••••••••••••••••••••••••••••••••••••••••••••••••••••••
This post resonates a lot with me and my approach to counter the idea that an LLM (as amazing as a tool it is) gives programmers permission to stop engineering. To me, executing a decision in the face of trade-offs is what the job is all about. So, I want to keep pulling the thread: is it worth reading—meaning, understanding—the code in order to prevent p95 latency from skyrocketing, avoid OOM death, and keep our apps reliable for customers? The cost of understanding the code is ostensibly very high relative to just using a clanker (citation needed), so does paying that cost translate to value—what is it worth? Concretely, if vibe coding increases bugs and enshittification, but decreases spending without decreasing revenue (read: customers suffer, but they don't leave), do we still need to pay the cost of reading code? Is it better to invest in stickiness, lobbying, and market capture? This comment shouldn't be read as advocating for this; it's a thought experiment about pragmatism and trade-offs, and reflects what I'm witnessing companies (and "programmers") asking themselves. I personally care a lot about understanding systems and code, and I believe I'm paid to do exactly that (I very much enjoy understanding code and nobody is paying me to run a business, so I may be biased in this). In everyday SaaS-land—not talking about critical safety systems—we've seen production databases destroyed, personal data leaked, platforms unwittingly exploited, and UX bugs creep into our operating systems, and yet the companies involved keep on keeping on. Is thinking , meaning taking the time to actually understand what our systems are doing, going to increase their shareholder value?
••
 Top