For me, it's like reading prose with "Not X, not Y, just Z": it's technically correct, but grates like fingernails on a chalkboard.
I have real trouble sometimes, reading what SOTA (Fable, etc) generate - no isolation or partitioning at all.
The worst was the planning an AI does. When I plan something, it'll be split according to data structures "An object to hold this, an intermediary for the obejct to talk to ORM, a serialiser for it that does this", etc.
The "plans" from SOTA are sometimes just hilarious. It'll go "phase one, implement these user-facing features. Phase two, implement those user-facing features", etc.
That's not a plan, it's an aspiration! A roadmap maybe. A plan, in my way of working, is a blueprint of where all the data goes, with algorithms connecting them. With AI, the data is incidental, the algorithms are incidental, only the goal (in the form of tests) remain. It'll work out some spur-of-the-moment idea around data at the time of writing.
So yeah, I do what you do and throw their stuff away. Currently having more success laying down a skeleton manually and then asking them to add a single feature at a time.
That's with SOTA models as of September-26-2026.
1) The size of task the AI has been given to do appears to be too big, which is why it looks like a roadmap/aspiration. You can ask it to implement a single feature or even a single part of a feature. Just keep cutting the size of the tasks until you become comfortable with it.
2) You like plans in a particular way following data structures/data etc. have you actually told the models this. It doesn't magically know. For the record the fact that the models focus on the end behaviour covered with tests is the way to to it imo. The actual implementation is less important and can be refactored as you wish fairly easily with the AI with the tests ensuring the feature still works.
I do agree though that the current SOTA models are very keen to just implement absolutely everything straight away without explaining/exploring properly. You can customise it fairly easily by using the various skills/agent/claude files to remember your preferred workflow, imo the agents adhere to theses better than they used to even just a few months ago.
since i really don't follow the idea of "writing unit tests first", i implement the feature, test as a user, and then use LLMLs to write the basic unit test. then, i will write more tests to make sure i'm covering everything.
feels like an ok-ish compromise because LLMs can do some ok job with defensive code, while i maintain the main implementation and more advanced test scenarios.
It usually takes me two or three iterations to get there though. Discussing design and principles before writing the bulk of the code is a must. And then a pass or two of review to weed out ugliness.
Still saves time compared to writing the code by hand. Especially for tricky things, where type checking and tests can verify correctness.
That's the whole problem with AI imo and I think the slot machine analogy is mostly right. It's just not predictable whatsoever and then you won't even be able to review all of the thousands of lines of code that you generate. You never know what you get and this has some serious safety implications that are not acceptable. Yes, it's fine as chat to just generate some snippets here and there that can be easily reviewed. Agentic coding is horrible imo.
Am I crazy, or hasn't it been this good for a very long time? The ability to get code out of it after correcting it, correcting it, specifying and respecifying, instrumenting and reinstrumenting, reviewing and demanding refactoring, "no not like that", etc. has been there (for me) nearly from the start. They're great when you're working on something you're not an expert at, and fine if you're working on something that you are pretty good at (if you like to have a cheerleader that sometimes trips and falls on her face.)
My problem is that they don't understand some things that are very clear, and after you've corrected them to get them on track, you're exhausted. You put all of those corrections into a file for them so that when they make the same mistake in the next session you won't have to wrack your brain correcting them, then they a) ignore the file, or b) make a bunch of spurious objections because they were all ready to object and the saved response killed all of the content of those objections. They still seem remarkably dumb.
They all still insist on serving bash code for FreeBSD where bash is not native. I've had to correct them all this whole year and they apologize profusely but, the next time, still use bash.
Note: you can install and run bash on FreeBSD. It's just not in the base installation or usage and I do not use it there.
There's no clear path moving forward. Overreliance on LLMs means your knowledge will exponentially decay and you will absolutely crash any future tech interviews becoming unemployable. Not using it for some quick wins feels wasteful. Finding balance between the two extremes in addition to all existing software development woes is really hard.
I have JIRA mcp wired up I tell agent to pick up the ticket it makes a feature. I review PR deploy to test server click happy flow through.
It works for me and company I work for. There is a world of difference between using Claude Code, Cursor configured properly and just using copy paste code with chat interface.
I am always baffled by how people take their experience as „this is ultimate truth”.
People are just recounting their experiences. It’s true my first experiences were all terrible with AI, I’m still just getting to a workable experience. It’s better when it can iterate a bit, that hides the fact that the first pass might have hallucinated junk that is smoothed over in subsequent passes.
He is showing no doubt backed by his extensive experience — while we have loads of other people writing about their experience that is totally opposite.
Well yes people writing about their opposite experience also have the same issue.
For me for example TDD never landed, it was always too much hassle for not enough return on investment. Scrum/Agile the oposite it landed quite well despite of what I read every day how useless it is. For me AI also landed, it wasn't that way 3 years ago, exactly 3 years ago I could have written the same experience as the parent poster.
When used in projects with good practices it was writing good code.
I've been finding AI either works great for people or is just awful, and I'm having a hard time pinpointing where the problem is. I don't think there's a big intelligence gap between us, so it's not like others wouldn't be noticing the things you're claiming when using AI. I genuinely believe your experience has been bad, but the part I'm trying to figure out is why it's been bad.
So my questions are: have you given things a fair shot? Have you tried turning things on its side to see if approaching AI usage a different way results in significantly better and more useful results? Do you actually put thought into what you're writing in your prompts (i.e., if I took your prompt and gave it to a junior dev, would they know what to do)? Are you just using the free models available online or do you have a dedicated integration into your development environment through Github Copilot or other vendor(s)?
I mean what I'm about to say with zero offense, but this honestly reminds me of how the elderly generation says, "technology doesn't work for me because it's always broken", and when you go and try to coach them on how to use the technology that's problematic for them, it's as if you're casting black magic.
I don't think AI is the second coming of Christ or that we're anywhere near AGI, but I do believe it's an extremely useful tool that people should be using.
That any of this would be needed is a problem on the AI side, not the user side. It can't be expected that everyone puts in hours of upfront work, trying out models and different ways of working just to see if AI is there yet for them, every time a new SOTA model comes out. Time is limited.
Also, the more thought, time, and effort needs to be put into prompts, the less useful AI is.
> I took your prompt and gave it to a junior dev, would they know what to do?
While overall I'm getting decent enough results, the AI often makes mistakes that no human would make, like not styling or aligning a new button on a webpage the same as the buttons right next to it. In general, it has major trouble with having a coherent vision of the entire project.
It's also obedient to a fault, never pushes back on anything I suggest. I'm pretty sure if I asked for a low user value feature that nevertheless would have to increase backend complexity 10x, it would go ahead and just code it, no questions asked.
> I don't think AI is the second coming of Christ or that we're anywhere near AGI, but I do believe it's an extremely useful tool that people should be using.
Ah, same. Just remember, we're very early into LLMs, they're still very much a sharp tool for experts, not really safe or convenient to use for everyone (yet?).
I mean, that's why it's not writing good code then, since if you're not using it with an actual harness then it can't read your current code and contribute. Also there is a huge gap between models made more than 1 year ago now versus today's models.
Just today it made three glaring mistakes in one session:
1. It read a file in the wrong directory, because that file had the same name as the file in the right directory. It apologized when I challenged it, promising me that it would remember to "read import statements" in the future.
2. It miscounted the number of times a function was called in my repo. It said 20, while my built-in IDE search accurately showed 17. Again, it apologized when I corrected it.
3. It referred to a variable by name that does not exist anywhere in my code. It apologized, and said it was referring to a variable used internally by one of the third-party packages installed in my repo.
So many apologies.
It's the little things like this that remind me on a regular basis just how little I can trust artificial "intelligence."
Without them, not only do agents get simple things like function call counts wrong, they tend to return different results. I use this as an example when showing people how the tooling works.
Grep is fine for simple use cases. A step up from that is ast-grep and I need to explore this tool more. But I had the most success building a small pipeline that reads the old code base, parses it file by file using tree sitter, and then loads it into a SQLite database for querying. For example, I have it capture construct definitions and usages and represent those as directed edges and nodes in a single table depending on the node type. The agent is instructed on how to query it and perform interesting queries like build call graphs, or determine dependencies between domains (modularity is not great in this codebase) which is helpful for us to extract around capability lines.
I also calculate fitness statistics, and have some code to capture specific details and knowledge about this very old framework that short circuits agent work in the future. We have some “interesting” magical libraries and functions that block static analyzers from going beyond the call site. This is mitigated, and means agents don’t have to “guess”.
Making all of this available to the different team members at my work has been pretty helpful. It’s faster (fewer tool calls), cheaper (fewer tokens), and accurate.
Otherwise, that’s exactly the tool to help those kind of queries IMO
Now everything is just like above, dorkspeak. Where it reminds me so much more of kids arguing in the playground about "who would win in a fight Darth Vader or Batman"; or console-vs-PC debates. Everything is about the genuine complexity of navigating certain products, of being first and foremost a consumer of something and putting all your energy into comparing various things you are free to choose from.
Its not even like its less techincal, or more mean now, or anything like that. It's just very different and I know its been a while but it feels like it happened overnight.
People can't even discuss about any politics anymore.
Back in Obama's days there were always interesting discussions in the political threads and they were rarely insta flagged. And it was mostly nuances wrt business needs and what that means for our societies. The PRISM news also frequently got heated, and still stayed somewhat technical.
Extremely noticable compared to the platform it is today.
Tbf though, HN always had a few topics it was extremely irrational about.
Eg Apple since the start and Musk post 2012...
As one of the users that moved to HN when Reddit closed access to third party apps it’s also quite literally this. I apologize for any reduction in quality of discourse I might have caused.
I'm amused by your surprise.
>Its not even like its less techincal
I'd wager that, actually, it is less technical. We constantly see otherwise very techinical people here poking their heads up and admitting that they haven't been writing any code for 6 months or more, and the ones that brag about it are seemingly unaware that they've been reduced to being a technical PM (the ones that contest this probably havent worked with capable technical PMs). No one wants to engage with vibe-coded Show HN entries because the poster may not grasp the implementation. No one appears to be doing (or sharing) anything super novel with all this coding superpower they have suddenly attained so there's nothing technically interesting to talk about. The last time I found an AI application submission interesting was the one where the person was trying to turn their pet's random keyboard typing into a language (if I remember correctly).
If the guidelines banned any AI comments that didn't preface their comment with the model and version of AI they used, their application domain, and the programming language being used I'm sure there would be less contentious dorky debate and more technical/practical discussion.
And since many here don't see the endgame with AI resulting in anything good for their career or society, like the endless remote work debates, they comment because from a strategic point of view they do not want to cede the narrative to the other side on such an important topic. understandably. hence more dorky debates.
> No one appears to be doing (or sharing) anything super novel with all this coding superpower they have suddenly attained so there's nothing technically interesting to talk about.
You don’t see a relation there?
i read some of these posts and it feels like the experience with boomers i had to help with their computers at my college job.
they did the least and expected the most.
Could you work out that it was wrong, given weeks to go and check its work? If so, your trust is misplaced.
You ARE taking days or weeks to go and check, yes?
I've never taken weeks to go and check bugfixes in the before times, I don't see why I'd expect it now. Once we know what the cause of the bug is, validating the fix and writing a test for it is usually trivial.
ive been regularly seeing this exact reaction online since November 2025, sadly :/
Working on more complex, logic heavy projects with strong performance needs I find the models to be useful but certainly not 'one shot' on pretty much anything. And often incredibly frustrating and genuinely bad code that collapses performance and bloats systems - like what a really bad junior might write.
When I'm working on large standard crud web projects with already decent architecture and a good harness and skills, honestly they work pretty well a lot of the times, the code looks good, does what it needs and fits in with the architectural style.
I really think a lot of hn people just write simple repetitive software, and a small portion works on complex, weird, dense, and novel'ish logic projects. The two obviously don't have the same experiences.
Yes, so I'll disclose it: I was using Opus 5.5, on medium effort, in Claude desktop, which has full access to my entire repo.
Now everyone can officially lambast me for "using the model wrong or using the wrong model," exactly as you say. But I find it quite interesting that one of the commenters here assumed I was using Opus 4.6, because these mistakes sound like that old version! I'm using the version released just four freaking days ago!
I expect some commenters will now say, "Oh, you should have been using Fable, you old boomer." To them I say: "Yeah, well my employer doesn't allow me to use Fable." And, in jest: "Now get off my lawn."
It's basically impossible at this point to take these people seriously anymore.
My model weights don't change unless I change them.
The most telling one is counting function invocations wrong, because that's simply not how models work anymore. They use terminal commands and Python scripts for research like that (if not an LSP, if one took the time to set up their tools most effectively).
Combined with their attitude, I have little doubt that the parent has disabled tool calls, is working in some janky Harness like chat/Duo/Juno, or is using a severely reduced or outdated model.
The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product or time spent maintaining and refactoring large code.
Anyways, if you want to continue on the path towards regaining control and take it to the next level I wrote something similar here: https://blog.sharefile.systems/be-brave-go-low/
Now I can just say "add 2FA" and in 5 minutes, while I test something else, it is done.
It also made iterations a lot faster, you can try something out, see how it feels, if it doesn't work, you can just trash all the code and start again.
Haven't typed a line of code or read any code for over 6 months now.
And I used to love coding and be a competitive programmer, but this is how "coding" goes nowdays.
I have a mental model of what it would do, and how it would work, and I ask questions to confirm things and tell it to watch for specific gotchas. Then simply test the feature myself a bit.
Security-wise, I think the latest cyber models are better than me anyway at finding vulnerabilitates and pentesting such features.
Plus 2FA is a very common pattern, so it likely has in the training dataset many really good implementations.
I don't doubt that, but they are equally good in making mistakes, over-engineering, or adding things you never asked for. They have all sorts of patterns in their training data from excellent to inadequate and I find them challenging to guide them consistently in one direction. Also with questions and tests, they can add something extra you didnt need and you dont know about, so your scrutinizing questions and test cases could miss that.
At least for myself, I didnt find them reliable enough yet to do what you describe and just not look at the code at all.
Models have found vulnerabilities that i wasnt aware of sure, but their fixes to the bugs they found often included "overengineering". In this case by "overengineering" i mean optimizing for passing test cases related to said vulnerability they found. eventually i have to step in to make things coherent and make sure that future agent can look at this part of my code and copy it to not introduce that particular class of vulnerability. Otherwise if i dont do that similar vulnerability and codesmell keep appearing throughtout the codebase.
I have increasingly automated encrypting and rotating secrets and setting permissions on them including better network level practices. Thanks to AI which helped me quickly implement those. So security wise i am better because of AI? But I also attribute it to my know how rather than the AI because I have never seen AI suggest robust but simple security postures.
I have no idea how the code looks like, and barely even tested the app entirely, because it is still not released yet, but I do think a lot about new features, tweaks, improvements, etc. Years of programming and game development did help, but I don't think anymore that code is relevant, as long as it looks ok and feels good.
[0]: https://ultimidi.com
>barely even tested
So the lowest stakes possible and you have absolutely no idea what bugs are waiting.
Also, I barely tested and kept changing things simply because of this: whatever I ask for, seems to work as expected.
It's one thing to test once and find 10 bugs, and another to test 10 times and find 1 bug.
This is worded as if knowing how the code works and testing it is hinged on it being released.
Even after release, I don't see reasons to check the code if everything works and people are happy with the app.
I basically only manually test e2e myself, other tests are automated, code review is automated. I can tell the model to test for me too specific things or to add tests for specific potential issues, performance benchmarks, compared different implementations, etc.
The focus is a lot more around the code than on the code.
1. If what you say is working, you have a working software factory that should be capable of matching the output of dozens of engineers.
What very impressive externally verifiable results have you had with this?
2. If you are working in software with plenty of customers, my strong suspicion is that there are people on your team who are looking at the code who furiously trying to reign in your output.
I think this one is really cool[0], will be a free piano learning app. I do have other projects, but they are all at around 80% too, because some systems are shared amongst the projects and have to be finalized too (i.e. now I'm implementing my own transactional/marketing email service on top of Amazon SES, I need it before releasing ultimidi so people can register and receive email confirmations).
If I had a working software factory like you describe, I’d expect you to have hundreds of apps of that level of complexity in a year.
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
But you don’t have any users much less paying users, so you have no idea if this system works when you do.
> I’d expect you to have hundreds of apps of that level of complexity in a year.
I am working at around 7-8 projects at the same time. The limit becomes me having to remember what I was doing for each one. AI can implement things nicely, but it really sucks at deciding which features and having its own ideas about novel game mechanics or UI/UX patterns.
I don't consider it a "software factory", just a more robust way to implement features, and it's still a WIP. One system I implemented locally is called "TaskHub", which receives some implementation or testing details from a SoTA mosel like Astra and implements it locally using Qwen 3.8 27b running on a RTX 3090.
I delegate things like app testing, navigating the app and taking screenshots, analyzing screenshots, etc.
And yes, this specific app doesn't have any users yet, I will launch it next week after I finalize the user authentication, payment flows and mobile apps. I will release it mostly as is, see user response, and then change it accordingly.
> I don't even know how "vibe-coding" works in a team environment, because for me now it feels like it's "ideas to app" directly, so dumping by brain/ideas directly into a functional product.
If you have a system that does “ideas to app directly”, that certainly sounds like what people are talking about when they say software factory. In fact my very large company would probably pay you millions if you could come in and make this work for our applications.
But more of, one feature idea, that the AI can add to an app.
So it's not prompting "a gamefied piano learning app", but adding a single new feature like "add a minigame, where the user can control a character on the musical staff, [... 20 paraphs later ... ], make a good implementation plan, implement and test the app until everything is production-ready and bug free /goal"
That is a single "idea" that also requires a lot of fine-tuning after, but rarely in terms of code, most of the times it's only in terms of functionality and design choices.
I notice this weird hostility whenever the topic of AI coding comes up and it's never made much sense to me. If someone told me about their new method for practicing guitar I'd feel like a real tool if I demanded they prove it for me then and there.
People don't owe you their "very impressive externally verifiable results" - /u/XCSme already posted their app in another comment, it looked fine to me.
You've made your ideological position very clear here, you don't need to keep heaping it on.
I can see why some folks here want to quibble, though. In other comments in this thread, they’ve compared to code quality favorably to the output of an average developer. That’s slightly insulting to the field in general (although, I guess most programmers have a low opinion of average code quality).
In fact it’s actually a very socially agreeable action, as it gives them the opportunity to show off their new skills without them looking like they’re bragging.
Now, if you happen to know for a fact that the guy cannot actually play guitar, then you’re just setting him up to embarrass himself, which is maybe an extreme punishment for the relatively minor crime of spouting some bullshit. I could buy that that is hostile, sure.
But you didn't write or read any of the tests so how do you know they are accurate?
You're acting as if code was incredibly secure before LLMs because humans were reviewing it.
People bash LLMs for overengineering but for this it's what you want. Taking extreme edge-cases into account that a human would never bother with and obsessing over security.
I mostly vibe coded a queuing system to replace something we’re using at work (last week. Spent about $1500). Then I meticulously went through the code.
It was much harder to review because it was ultra defensive and included guards for tons of edge cases that weren’t possible.
Unnecessary abstractions for possible extension later. Useless indirection. Probably 3x as much code as there would have been if I’d written it by hand.
I didn’t one shot this. I kept a pretty tight leash on the AI. I had probably a dozen markdown files with of plans that I created over hours of back and forth with the AI and reviewed before each implementation round. I had automated reviews and quality gates etc…
What I found in review was that it was full of very subtle bugs that would have bitten hard in prod. Committing offsets asynchronously that would lead to dropped messages. Clock drift bugs that would lead to dropped messages or write amplification storms. Lack of back pressure in some stages of the pipeline that would cause notes to get silently OOM killed. Weird over-insistence on never crashing in most places that would mask systemic errors.
If I’d just shipped it without review, it would have mostly worked. But at the scale it’s going to be used (tens of thousands of messages per second) it would have caused production issues for months while we tracked down each of these issues.
It's not possible until it is. This is the justification lazy developers like we all are have been using leading to bugs down the road. This glorification of hand-made code is strange, like we weren't writing dirty code full of shortcuts and hacks all the time.
Overly defensive code is harder to read and change for both humans and LLMs.
And many times it makes debugging harder by moving or suppressing failures.
> I ask questions to confirm things
Oh my.
As someone who reads the code, I can tell you, asking questions to confirm things is inadequate. The models lie to me, daily.
Every day I have two experiences:
1. I’m blown away by what it can do
2. I say, ”wait, you said this, but the code shows that, so you were just going to leave that endpoint without requiring any authentication??” and I get the “you’re absolutely right, that was my mistake, and you’re right to call it out” song and dance. Daily.
It also adds all kinds of bloat to code, tests, and “documentation”. I’d say I spend ~30% of my dev time picking lines of code or documentation and asking, “why does this exist?” and “what would break if we deleted this line?” and then arguing with it and removing things.
Truth is, modern software was already quite shit and full of bugs. All major apps had bugs, issues, going down, etc, so users did get used to things not working. I honestly beleive AI coding nowadays, for better or worse, does things better than the average developer.
Yes, it is overly defensive and verbose, but the end result is in general ok and fully functional. Yes, it adds 30 tests and "release gates", and they are not even that useful, most of the times they just act as an extra safety mechanism to make parts of the code immutable, so release fails if the model accidentally changed things.
Another issue with looking at code, is that it's very hard to manually change things anyway. I can't just change a variable from 10 to 20, because I don't know where it is used. I have to ask the model to set that value to 20. It is quite stupid and inefficient, but this is one cost of coding using AI. But, if you do this, things will likely work.
That being said, I've mostly used Astra xhigh since it was released and things just work.
This should be a giant flashing red light. If you can't figure this out, either 1) you're too junior to be effective using AI, or 2) the AI is doing a truly awful job organizing the codebase. In either case, it's a sign to slow things down and understand what's happening before proceeding.
Is that that variable might be used in tests, UI, docs, agenr markdown files, other related projects too.
The design spec could say "always use 10px margin", then in code we have a const with value 10 used everywhere. If we update 10 manually in code to 20, then the design spec is now outdated.
You can't have it both ways. "The AI always does it right" and "I can't keep track of all the places where the change must be made, but the AI can" are incompatible and contradictory ideas. If you can't keep track of all the places that need to be changed, you can't check that the AI did it right.
I dread to think what that means in security conscious code.
Can you give an example of unsafe defaults used?
EDIT: Come on? Won't post your code for everyone to see? Why not just put it all in a repo, client and server both?
EDIT 2: Amazing. If you punch in notes on the keyboard for like 10 seconds then click the keyboard-icon button on the bottom right the website crashes
EDIT 3: If you click the main CTA then click "Let's start" the website hangs and you need to manually refresh the page for the content to load
The game seems quite bug free though, including minigames. The UI could be better, but it's not done yet.
I do for example have an automated system that simulates progression, takes screenshots of the game to find potential hidden buttons or overlapping elements, to test for performance, etc.
If it looks like a duck, and quacks like a duck, I honestly don't see why I would review 100k's of lines of code.
EDIT: I might have replied in a wrong thread, but it was about this entirely "vibe-coded" app: https://game.ultimidi.com
> why I would review 100k's of lines of code.
If the core or that app is more than a few thousand lines of code, something is seriously wrong.
I don’t want to shit on your app. It’s cool. I’m glad you built it. I’ve vibe coded all kinds of toy apps for myself and my kids.
But it’s not strong evidence that code is irrelevant.
> If the core or that app is more than a few thousand lines of code, something is seriously wrong.
That's the thing about vibe-coding: it is not the core, it is the entire app. We no longer make MVPs and release those, with AI we can make directly the app including all bells and whistles, entire progression, not just one level, all the systems around it.
Why? Because if something needs changing, it's just one prompt away. I do think code is fluid now, any choice of architecture can be instantly changed at basically no cost.
Maybe my mind is just finding ways to cope, thinking that I "wasted" thousands of solving coding challenges and fixing bugs, but I do think, for better or worse, that manually coding is gone. Same as we no longer code in assembly anymore. We no longer write C. We no longer write JavaScript. We no longer write TypeScript. Maybe not today, people don't like change, but manually writing or even viewing code will only be done in a few educational and high-performance/risk cases.
Come back to me when your app has users and adding new features subtly (or not so subtly) breaks every work flow that you haven’t explicitly tested.
You can’t commit the prompt and regenerate the source code each time because the whole reason that an LLM is useful is that it makes thousands of decisions for you. And those decisions are different each time you regenerate.
The only way to enforce that those decisions are the same each time you regenerate is to encode all of them in tests. But at even moderate complexity that leads to an overconstraint problem that will halt development.
We see this when using LLMs on large apps. Anthropic gave up on their C compiler. Even with an unlimited budget they stopped being able to make progress on it.
I see this in some games I made for my 4 year old. I had them one shot some “juice” when he gets an addition problem right. Combination of screen shake, sounds, flashing light, explosions etc…
It looks pretty cool, but when I tried to tweak the animations with prompts it was always worse. I eventually went in and edited the code myself and I could see why it was so hard for the LLM to change anything because it was a horrific mess of interwoven animations.
I was able to pull everything apart and manually adjust what I wanted.
But yeah, overall you have to be ok with the app being approximate too. Maybe after an update a button is a different size, or in a different place, or it suddenly has an animation to it. Those smaller things are a bit harder to control when making changes at scale, and if not clearly documented.
For me this is not necessarily a big drawback, for things like games, the core game loop, performance and game feel are a lot more important than any small UI tweaks.
Hopefully, the better the models get, those side-effects will only be improvements, not degradations.
Now come back to me when you have paying users.
Better yet come back when you have paying users who depend on your app to do their job. And in addition to buttons changing location, you are constantly breaking their work flows because they are using the app in ways you didn’t anticipate.
I only started using AI for development for this product (13 years developed without AI) a few months ago, and customers are really happy with the changes.
I managed to implement feature requests that were pending for years. It took probably 1 month to implement what would have taken 1 year without AI.
>It took probably 1 month to implement what would have taken 1 year without AI.
1. I’m not saying that AI can’t speed you up. I’m only saying that you can’t ignore the code for anything beyond a toy app (at least not sustainably).
2. How much is that is down to motivation though? You’ve been developing something for 15 then years then suddenly there’s a brand new development methodology that is fun to mess around with and addictive.
I wrote about some of the issues I uncovered auditing my own vibe coded project in another thread https://news.ycombinator.com/item?id=49856310
I thought the same, but I do disagree with this now. Future software development will be mostly just creating black boxes and describing what the box should do, without ever caring what is inside the box. For now, we still have to guide the AI and tell what architecture it should likely use, or which libraries should use (just for the sake of performance and ease of development), but in the future this probably won't matter either. I honestly believe code does not matter anymore. What matters is knowing what to test when building an app, how to define performance metrics and knowing what is good/possible when developing an app. I know I won't be able to create a shooting game that supports 1 million CCU on a single vCPU. But I know that the input latency should feel good, and game should run smoothly at start and over time, to tell it to implement tests to check for memory leaks and avoiding JavaScript GC pauses, use object pooling when possible, etc. The complexity moves from telling how/what to code, to knowing exactly to tell it how things should behave and what's a good outcome. If I tell it "make sure bullets are object pooled", it will likely implement it properly, as object pooling is a very common pattern that exists in its training data, and it usually either works or doesn't, when it doesn't work there are obvious issues like objects shown at the wrong positions or not spawning properly, so the issues with the code would be reflected in e2e testing anyway.
2. That was both motivating and demotivating to be honest. The app was my "baby", having spent a lot of time designing everything, optimizing, carefully choosing libraries and make cool implementation decisions. Now I feel like all the newly added features are not really mine, and it doesn't even feel like my product anymore that I can proudly say: "hey, I wrote all the code for this app". It is a really demoralizing feeling, but at the same time, I like how all my ideas can now be materialized. And it actually works. And it works well. I am still getting used to the fact that I won't have full control or understanding of how the code works, and it pains me that this is the case, but there is no way I could achieve better results by manually coding. I would rather have a feature having 90% of the ideal possible performance and UX, than not having that feature at all. Plus, that 10% is still doable, it just requires a bit of testing and asking the AI. The problem is that most of the times that effort is not really worth it, not for me, not for the clients. There are a lot of other low-hanging fruits that must be addressed, and that's how software development always worked. Now I am happy that I can actually do nice things that before I would have never spent the time on, like making sure a specific settings menu has better UX on mobile (before I would have probably just made an element smaller to fit mobile, for ok usability, but now I can tell it to design an entire new UI tailored to mobile for that specific feature).
It is definitely addictive, as it comes with instant gratification, as opposed to slowly coding and spending hours before you see any results.
I agree with your linked post, and that is sort-of a big issue (even though, most of the times, those guards are not there for the functionality itself, but for the AI, so that in case things break, it gets a more clear error of what went wrong and it knows how to fix it better). This is my entire system prompt, which I assume fixes some of that defensiveness issue:
# Engineering style
- Prefer simple, readable data flows and strong invariants over layers of defensive checks, fallback branches, assertions, and recovery mechanisms.
- Validate at real trust boundaries, then let well-typed internal code rely on those validated contracts. Fix the source of invalid state instead of spreading null checks and guards through consumers.
- Keep code concise and add useful comments that explain intent, ownership, or non-obvious constraints. Do not add speculative protection, tests, or abstractions without a demonstrated failure mode or requirement.
- Use ASD-STE100 Simplified Technical English when you talk to me or caveman-like concise speech, be really directI have a practically unlimited AI budget and I use it all day long. We have entire teams of very smart people trying to make it possible for PMs to turn jira tickets into features.
We ain’t there yet.
My strong suspicion is that we won’t get there without something close to AGI. And if we get there, there won’t be any white collar jobs left much less software developer jobs.
If AI can take your raw idea and turn it into code with no input from you, it will be able to generate the ideas too.
> I agree with your linked post, and that is sort-of a big issue
This was last week with frontier models. Practically what it means is that I need to audit the code. Approximate changes don’t work when you move past very small scale projects.
And with epsilon big enough any output it good output...
Also, what do you mean punch in notes? Like mashing keys and pressing 10 buttons at once?
Never had the app/website crash, tested only on Brave (desktop and mobile) so far.
EDIT: I can reproduce the crashes, thanks, will tell the AI to test on Firefox too, testing on Firefox mobile specifically might be trickier to debug actually.
I have spent two weeks using opus just to write a plan/design for 2FA and iron it out until review (about 7 of them) doesn't flag it with 20+ problems (with security holes of various sizes), for which I had to guide it through to not turn it into a mess and whac-a-mole.
It's just a secret key that an autheticator uses to generate a time-based code, which the app can validate before completing a normal log-in flow.
What model did you use?
Astra xhigh on fast mode can probably indeed one-shot that in 5 minutes.
Plus optionally QR image generator to easily add that key to the authenticator app.
There are already many libraries doing 2FA, but implementing it in any language is quite trivial, right
No?
Aside from the fact that the implementation must be secure, you want for example to:
- handle accounts that have lost their second factor in an way appropriate for your business - decide what to do with accounts who don't configure it. If e.g. you want to send them authentication codes via email or SMS that's another can of worms.
Let alone the simple things such as making sure that your implementation works with the various TOTP apps
One-time displayed recovery codes are a standard practice, and most good LLMs will add it by default without even asking for it.
And yeah, I am talking about TOTP apps, I think sms/email is not as secure.
I agree that with a clean codebase it will be simpler, maybe couple of days, just running code review workflow takes 15-30 minutes and then decisions llm makes for each problem is often not good and lead it into overengineering rabbit hole, which means I have to think about each problem and prevent it from escalating.
But yeah if "look, it sends the code and I can enter it to login" is enough validation then it can be made in 5 minutes, sure.
The critical part is providing it a way to validate its work end to end. Without that, it's similar to asking a human to implement a feature with pen and paper.
Yes! Part of implementing it "one-shot" is actually a loop of planning, implementing, testing e2e on various devices and various edge-cases. It's not just writing the code, and this is not how "vibe-coding" works nowadays. It's not just writing code anymore, this is why most providers worked on their computer-use support too, not only for doing tasks, but also for being able to check and test the work they do e2e.
EDIT: Support was not the case, because the app is self-hosted, but the docs should indeed provide instructions for recovery if both authenticator and recovery codes are lost.
I believe "writing the code" means literally just "writing the code" not thinking how to implement it.
If the problem and implementation is so well defined, and "determnistic", this means that LLMs should also be able to just "write code" from specs, without thinking.
Removing the opportunity cost doesn't eliminate the other two costs of a feature
I'd expect it's in the tens of thousands, given how effortless you find it.
Code is a liability.
https://github.com/prettydiff/aphorio/commit/0c730389c9e4e69...
It always only took a few hours. And yeah, whatever, sure it adds up. But now you've got this dogshit looking flyer outside your restaurant, and it would have taken you like 45 minutes to adept a Canva template.
For AI, the skill loss and all the rest of it, the lack of control, the lack of knowledge of the codebase. It doesn't feel like saving the few hours is ever worth it.
As opposed to writing a prompt for 30s, then doing something else for 15m while the AI works on it? :)
> The real challenge is forming the new ideas in the first place and most of those new ideas coming either from using the code as a product [...]
Exactly.
AI is absolutely excellent at rapid prototyping. Prompt -> result -> use the product -> prompt. Repeat until the right design emerges. Once that happens, the code can be cleaned up, or even rewritten from the grounds up, taking into account what was learned.
This is using AI for productivity in any domain, in a nutshell. I just wrote a book using Claude as an experiment, and while the thing got done and it was an amazing tool and a great experience, what I’m left with is a book where every line needs rewriting, there are logical inconsistencies throughout, and the style is so bad it should actually just be binned rather than rewritten.
How was this a great experience if what was produced needed such extensive changes that your own assessment is that it should be thrown out? At what point is the necessary rework so much that the thing being reworked didn't really contribute much to the end product at all?
Getting a good writing style out of them requires careful prompting and many corrections, their default writing style(s) are so highly reinforced by training that they will always tend to drift back to them. Maintaining continuity requires you to create a lot of documentation outside of the text itself. It's a much more manual process than working on a codebase where you've set up a lot of automation and tooling that allows them to check their own work.
Edit: it's also worth noting that many LLMs have gotten much worse at writing prose as they have gotten better at writing code.
If you enjoy the craft and the creative process of actually coming up with new thoughts, instead of relying on the probabilistic combinations of thoughts of others, then you can just as well do it yourself and have full control of the process.
I use AI for low value work with dead lines, where the customers don't really care about the result either. For the golden services and customer engagements were I can tell the customer cares deeply, I use little to no AI, and then get a deep sense of fulfillment due to a job well done.
In effect, everything is classified as: "low value work with dead lines, where the customers don't really care about the result either."
Maybe robot arms can solve this last part one day.
If used naively, where AI spits out 1000s of lines of code, that is far from perfect, written in a style that might not be what you are used to, it can take longer to parse, than if your colleague of 2 years wrote it.
Have you noticed it, too?
And just like that, the book got abandoned.
N.b. some of its analysis and laying out of faults in arguments was actually pellucid and brilliant, it can’t be denied. Just it comes with prose that can’t really be used for anything. And even on another occasion when I got it to help me redraft and extend a different book of mine, then it randomly and consistently started stripping out all the stylistic flourishes out of my sentences, to the point where it couldn’t notice that word choices were very deliberate and actually set up little punchlines and logical payoffs paragraphs or chapters hence. And even when I explained and showed it what it was doing, it was like “ahh that’s so clever and brilliant” but just continued to do the same thing.
I think this is a common experience, no matter what field or era you're in. Once you start applying QA to some promising new process or technology for the first time, the whole cycle slows right down.
But that's actually the point where things get interesting, and innovators get to roll up their sleeves. I wouldn't give up!
Playing an instrument vs electronic/computer music.
We do forget skills we don't practice, especially fine motor skills (like playing the guitar or typing code).
There's inherent pleasure in playing a musical instrument - practicing improves fine motor skills and produces satisfaction.
You can play for yourself and that can be a great experience.
Often people create music for other listeners - and now the satisfaction comes not just from your skill, but from how the music impacts your listeners.
They say you can put more of your 'soul' into music made with an instrument, but I'd say there's quite a bit of electronic music with just as much soul.
People who create electronic music don't generate any of those sounds with their fine motor skills, but they do have a plan about how the song progresses and what emotional state it elicits in users.
That's why you have DJs which are more popular than others.
If you stop playing the guitar for a year, then pick it up and try playing something, you will feel very rusty. But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.
I guess my point is - If you go full on agentic, you'll loose some of your coding skill, but you can get it back fairly quickly if you go back to manual coding. On the flip side, you get better at using AI if you use it, so your thinking is at a higher level, but you give up understanding the low level details of how exactly the code works.
Either way you're making 'music', albeit a different kind of music.
90% of Hacker News posers^Wprogrammers
Now, manual labor does have its place as a form of art--take high-precision hand-built timepieces for example.
To add some points on he other side of this analogy:
There is not a lot of purely electronic music that has stood the test of time, at least not when it comes to popularity or, more relevant to the metaphor, profitability. There is usually at very least a human voice in the (literal) mix, but more often than not there are also traditional instruments mixed in.
Take this next point as you will as I am being a bit tongue-in-cheek: While making music-making more accessible to more people is totally great, if I could go a week without hearing a variation of the phrase "Check out my dark ambient drone project!" I would feel oddly accomplished.
Most importantly, though:
> But give it a week of practice and most of your skill comes back.. and in 1 month you're back to your peak skill.
This is only true if you had the skill to begin with. For many electronic musicians, by which I mean junior developers, this is not the case. Does it matter? As a 45-year-old traditional musician... er, I mean hand-coder... I think so, but also ¯\_(ツ)_/¯
No, playing an instrument VS making electronic music has absolutely no comparison to writing code by hand or with AI.
You're comparing the difference between a motor skill and a knowledge-based competency, with the difference between two knowledge-based competencies.
Of course if someone wants to stop using AI completely that's a completely valid decision[0], but I somewhat feel like AI is just a tool that can be easily misused.
I constantly have to review giant PRs and I noticed that I'm handwaving them more and more often. We went from almost no commit messages to walls of text that no one reads. We're starting to become bottlenecked on reviews because code is coming out too fast.
But at the same time, these are mostly issues stemming from a lack of understanding of why some of the standards/processes existed in the first place. If a developer thinks the commits have to be written just to tick a checkbox, they won't care about making them readable.
And at the same time, I'm getting a lot of value from AI, in tasks that do not necessarily have such adverse effects:
- I can create quick tools to test something, or parse/process some data. In these instances code quality is not important and I don't really want to spend hours on developing it myself (just to feel accomplished?)
- I can research issues in our codebase by just providing a log file. It's not always gonna be accurate or correct but it often gives me a very good starting point, almost always quicker than I could've done it myself
- While I do not use AI to completely generate ticket descriptions, asking it to generate me a body containing the relevant code snippets and references allows me to focus on verifying that what I'm writing is correct and understandable.
Etc etc.
So I don't know if it's just the nature of my work, the fact that I have a different skillset, or different priorities. But it somehow feels weird to me wanting to completely abandon AI just because in some cases it can lead to frustrating consequences.
[0]: I too just started a new project where I'm forcing myself to use absolutely no AI!
I use it to find reasoning gaps, add examples, add citations etc. The LLM/Agent can find them quicker than I.
A week later, same prompt, really low poly blender model. Either they reduced token usage per person, or they quantized the model, idk, but it really doesn’t work as good as day of release anymore
It all came back every time I went back to the trenches. The question in my mind is, what do "AI-native" engineers have that they can come back to?
- During these decades, I have relinquished control several times in favor of productivity. From knowing exactly where every byte is placed in RAM and where each cycle goes, to only knowing that for the inner loops, to just knowing the machine code that the C compiler will generate, to dynamic memory and classes and indirections and cache misses, to wasteful but oh so very expressive javascript and python. I stopped writing my own engines and used Unity, Unreal Engine, Phaser, Godot...
Relinquishing control is easy if you are still truly in control of the new layer, and know where the pitfalls are. Where are new engineers going to gain that expertise?
- Regarding addiction, I've also quit smoking. After 25 years of daily cigarettes, one day 16 years ago I just stopped and never touched another one.
I won't pretend that applies to everyone, or even that I'm impervious to other addictions just because that one was so easy to shed. But harder or easier, everyone can stop problematic habits if they are clearly problematic.
- I don't know where we're going with all this AI. I would prefer it had not happened the way it is happening (IP theft, job destruction, race to the bottom, power concentration, etc). I love progress but I don't think the most important aspect of progress is how fast it happens. Speed only helps the greedy and the terminally ill.
But I'm not going to pretend it hasn't, or risk whatever is left of my professional future boycotting it in favor of a different reality, or (who knows) reject a medical treatment just because it was proposed by Opus 7. The world will live or die regardless what I do, but MY world relies on me.
I will continue trying to have enough expertise, passion and attention to detail in what I do and how I do it, that whatever level of control I have over it is as optimal as I can. From typing z80 bytes, to asking Claude to change a 5 for a 6, the above traits are what has always mattered in my experience.
I see great engineers troubleshoot everything by pasting logs into the prompt and blindly accepting the answer. Zero added value while they ctrl-c ctrl-v themselves out of a job.
For the rest I do not. I do not place AI-generated code anywhere.
You lose all control AND UNDERSTANDING.
When things go wrong it gets very messy.
I will keep doing this, I think it works well, I emjoy programming and I think it is productive.
For testimg I tend to write randomized testing, which takes a bit of design but oncr you have it, well, it os test-generatove and increases the quality of checks.
"See?? Here's that DEAD BEEF CAFFEE again! Look! Again! The FECE FACCA AFFEC7!! I'm so close to crackin' it! Aha.. Aha.. ABEBE23.. BECACA17.. 1337C0C.. It all clicks in place, don't you see? I'm totally getting it!"
He was all bubbling like this throughout the whole night until his brain just issued a shutdown to let the body rest a bit. That was truly a horrible sight.
I remember him every time I see instances of AI psychosis around.
Smh new conspiracy theories every day. “Ai is actually a drug and you only feel that it helps you but it doesn’t”
You’d rather paint AI as if it were a hard drug to cope with the world changing around you. I mean listen to yourself. Im not the one making conspiracy theories.
Like toddlers!
Soon we rediscovered Little’s Law. WIP was piling up and we were getting overwhelmed at the integration phase, and realized that we had got really good at starting projects but actually finishing them was a struggle. Tickets were moving fine, of course. Our rate of generating code and committing PRs was through the roof. But getting actual projects to a point where the stakeholders and customers were happy with the result was just not happening.
So now we have gone back to strict WIP limits and requiring every non-trivial project to have at least two people collaborating on it. The rate at which we are churning out code has gone back down, along with the token bill, but the logjam is clearing. Better yet, the stakeholders, who never cared about our quantitative velocity metrics in the first place, have eased off on complaining that we aren’t getting anything done.
I have not lost control.
I my most prolific project I do not review the code, but I QA test extensively.
In other projects at work, I review the code.
I prompt to simplify, I challenge implementation that solves irrelevant edge cases, resulting in much smaller PRs.
In projects where I do not work alone, I still write two line PR descriptions myself.
Dumping paragraphs of AI output into the description of a MR where I ask others to review I consider disrespectful.
---
> If you turn off your brain, and relax babysitting AIs, you’re not getting any better. You’re losing value
I'm hardly turning off my brain here.
As the author notes, the context switching and so on takes concentration and effort too.
I can say without doubt that I am more productive than ever.
I am getting better by the month, and I am not currently losing value, until the AI fully replaces both me and the author.
This is a personal time management problem. Not a tool problem.
THere is a lot of temptation to do many things at the same time. For some people this works. For others it doesn't. This is not a tool problem. You simply need to fit the tools in a way that matches your personal workflow.
I can do at most two things at the same time. A friend reads a book when Claude works on a task. We're all different.
Yep. Part of the job of a software developer is managing complexity.
It's critical because of its exponential nature: as complexity increases it becomes exponentially more difficult to successfully change, understand, or support the system.
From any given position, AI seems to be the right next single move. When things are simple, it's fast and smooth. When things are complex, sota models can handle it, accounting for the details of dozens or hundreds of things at a time.
The problem is, without strict guidance, it tends to add complexity. If you let it, it builds systems at a rate and in a manner that a human cannot understand it. The best next single move leads to a poor overall position.
Then you have no choice but to keep using AI, and to keep upgrading to the next, more powerful model.
Or retreat.
It's pretty clear to me that, as a whole, we don't know yet how to develop with AI. People are mistaking token output or short-term progress for real progress. They people who are going to be successful with AI-assisted development are going to be the ones who can maintain control of their systems while they do it, and the key skill will be complexity management.
The barrier to software development has only ever been computer access and knowledge.
With AI, it’s roughly computer and internet access.
This means we’re getting a lot of people who aren’t good at either software development or AI automation playing with both. It’s the majority of what people seem to talk about.
I don’t think this is bad, but I do think it’s making real progress in AI automated software development on teams which are good at both much less visible.
A conservative team member of mine estimated we’re working at 200x speed these days, compared to 2 years ago. And we still see ways we can improve. A parallel team is only seeing an 1.2x increase, but they are unable to modify their architecture around AI.
Some of this is shifting roles. You can have a mildly technical domain expert vibe code the frontend for a new module. The more AI automation you’ve architected for, the faster they can go and the higher quality the outcome. We’re experimenting with mixing vibe coding with specifying formal requirements to push this further.
This works well. And now you’ve cut dozens of rounds of the PM not knowing the right shape for the new software out of the process. Even if we threw the end code away, this would save us tons of time.
This is just one example.
If you cannot read it as the author, what hope do I have to read and make sense of the wall of text which doesn’t seem to describe what I actually need to start reviewing.
I really really encourage everyone to write their own descriptions for PRs. If you cannot succinctly describe it in a way another human understands then you don’t understand your own change and you should withdraw your request.
I think there might be (dare I say) a middle ground to get the productivity of the llm, esp as we evolve them, while still maintain a global and even fine-grain comprehension of a code base.
It is not a simple change, however, but a fundamental one.
Overall I think we are still living in the past and try to apply ourselves to the future. But if the ai craze is to be taken clear-headedly for what it is, it is a complete break from the von Neumann computer and all its resulting artifacts. So why should we use the same tools?
T3 Skynet already exists and is hunting John Connor down
I wanna see the conference room meeting where they decide to push an unfinished, unstoppable technology. I guess T1 is the closest as it happens when Skynet "gains intelligence"
> I wanna see the conference room meeting where they decide to push an unfinished, unstoppable technology. I guess T1 is the closest as it happens when Skynet "gains intelligence"
This is the mid-point of T3. From the Wikipedia plot summary:
General Brewster is supervising the development of Skynet for Cyber Research Systems (CRS), an autonomous weapons developer. The Chairman of the Joint Chiefs of Staff pressures him to activate Skynet to stop an anomalous computer virus from invading servers worldwide. Brewster fails to discover that the virus was Skynet becoming sentient. John, Kate, and the Terminator arrive too late to stop him from activating it. The T-X appears, fatally injures Brewster, and controls weaponized CRS T-1 and HK drones to kill other employees.it's people (Sarah and her son John who are against the coming skynet machines) fighting people who think the machines will be a net positive (the scientist making machines) AND
it's people (Sarah and her son John and the scientist they convinced the machines are bad) fighting people who think the machines will be a net positive (the company of the scientist which they break into to destroy)
It baffles me that someone who prefers TDD can't just stick to writing the tests and then using them to validate and constrain the AI's output. Only making a small number of changes per PR is easy to enforce as well. The problem here doesn't seem to be using AI, it seems to be a lack of discipline in how you use AI.
>There were tasks I could have done in 20 minutes easily, that took 5 minutes of an AI agent, and then 2 days for me to review.
This shouldn't happen at all if you constrain the problem. I've seen it happen many times when you just YOLO a vague spec and just keep prodding it to continue without paying attention to what is being done. If you have a well defined task that should only take 20 minutes of your time then it should only take 10 minutes to review the small amount of output. If the PR is 1000 lines changed then something went wrong, throw it away and rewrite your specification.
I think OPs point though is that, like addiction, discipline is hard to maintain. Humans are irrational, and sometimes, some people just need to go cold turkey.
productive people who feel alienated by AI and want to just opt out.
Just having that mental paradigm will lead to more correct uses, and avoid a lot of the most ones most likely to lead one down the primrose path. This cultural moment of SF AI psychosis will pass, but we'll be stuck with the resulting code for a long time.
Also, +1 to the folks saying that code is a liability. More is almost never better. Aim for succinctness and brevity. That cuts directly against the thrust of LLMs' tendency, but will give you plenty to curate.
On the one hand, a majority of developers oppose mass data collection and privacy-invasive technologies like ad tech. Many of those same developers also oppose crypto projects because of their energy consumption and environmental impact.
And yet, those same developers have embraced frontier language models--technologies that have simultaneously scraped vast amounts of the world's data, raise serious privacy concerns, and consume massive amounts of energy.
Tbh, I felt anxious doing it wondering if I am still able.
This means you need more practice to master the tools.
I'm not saying every PR will be perfect but if you get unwanted slop and you care deeply about that then you need to up your prompting game.
"I see luddites"
Resistance is futile, you will never code faster than AI, with less bugs, more optimized, with more features, in 200 languages, for a dozen of platforms, desktop, mobile, web, responsive, embedded, a thousand times cheaper than you, in your invisible niche market share (they already found you), not gonna happen, and then you'll cry in a corner that you were laid off, or your business got steamrolled by a new competitor selling slop that nobody understands and nobody will fix either, and you will never accept how people fall for this delusional mania if the beauty of art is in being hand made character by character in a punch card
Slop yourself or get left behind
We’re using an online, undeterministic, black-box middleman to generate our code. It’s 100% Trust me bro. No proof, no scrutiny, no guarantees.
> let me tell you about this experience, and how it was turning me dumber, lazy, and a worse developer.
Though the article discusses from the point of using agents, I digress to the topic of building with AI in general.
My experience has been the exact opposite. A new idea (usually related to correctness or architecture) is discussed first with the LLM where it defaults to average Joe idiotic bullshit pushback.
This frustrates me and I abuse the LLM for being idiotic by explaining the how. This results in a more refined and concrete form of the abstraction leading me to even more insights.
The LLM remains an idiot. But a useful idiot nonetheless.
A machine that does not learn from the feedback. The only thing these LLMs know about after the initial training is what makes it into the context. Imagine if someone couldn't learn anything new without someone else explicitly telling it what was important to the task every single time? You wouldn't call them a junior engineer.
I remember my parents saying "be careful with the bad guys who mow your lawn, fix your appliances, and drive all you kids to school. those are not people who are helping your family get things done more efficiently, those are drugs"
seriously what are we doing with this "AI is a drug" metaphor? Drugs aren't reading logs for me and writing unit tests?
> Once you start using AI coding agents, things spiral out of control very quickly.
no? maybe AI is not for you?
> Except you don’t write the code anymore, you just ask AI to do it, and half of the purpose of TDD (not biasing the tests by how you’ve implemented the code) is gone. But you feel you go so fast that you start not caring.
do you use code review tools? Did someone tell you to stop using them? Review your LLM's code, leave comments, tell your LLM to address the comments. Have you ever worked in management /architecture? Tech managers do this all day long before we had LLMs, it's not new. Except your workers are the smartest junior programmers you've ever had (and that is where AI is a problem, for sure - I worry for junior programmers today).
> It would not be so bad if things had stopped there, but that’s not how most human brains work. If you like something and you can have 2x, you’ll have it. I started pasting the whole Jira description of a ticket, and let AI implement it for me. Yay!! So powerful!!
totally! Have you not already implemented 10000 issues on your own and it's not boring yet? time for issue 10001 then? Whatever floats your boat...
> Control is an illusion. Other folks I’ve discussed this with agree that they don’t know 100% of what the code they are pushing to production actually does. I bet we not even 20%. Fucking scary.
sure have you ever accepted PRs from other people? ones that took a lot of work to read and understand? How is that different? except their code would actually break all the time if you didnt read it and now it's your bug to deal with because they're gone.
> Because if it was just about approving PRs that are perfectly written, it would be good. But for each PR I had to look at the code, the tests, the style, give feedback, switch to something else, go back again, push, see if I have any PR comments,
How do you think large projects get built? all by just one person so nobody has to review anyone else's code? (and even if you are - you still should be putting your own code up for review, reviewing it, and giving yourself comments)
> if CI passed, wait, now the linter complains.
why isn't the linting part of your CI ?
> The description is for others. Because there has to be a description, right? AI is so fucking good as sounding professional, that you relax and let it be. It’s not that I actually thought the code was good, but you get convinced over time, you get lazy, you become complacent. You stop questioning, and start accepting as good some code you would have never accepted, just because you cannot tell why it’s bad. You have lost control.
Just seems like a lot of issues here, so sure, dont use AI it's clearly not your thing. i do not let one line of code I dont think is "good" get committed, period
> Most AI developers will deny it, but deep inside, they know it’s true. Just don’t want to face it.
but dont assign my feelings. speak for yourself. I'm doing this shit for 40 years and you might find it gets a little tedious and repetitive after awhile
> Say no to drugs. Kind of a metaphor, but not quite.
it's a 100% totally wrong metaphor and it needs to die
I generally don't use it for code - I've been writing code for 30 years, so I'm pretty quick, and it's more productive when I look at total time for me to just write things myself.
But not always.
For example, I was recently writing some firmware in MicroPython and needed to implement Bluetooth to control the device.
It's been years since I've written python, since before asyncio. I'd never written MicroPython and didn't really know anything about Bluetooth or BLE.
The AI taught me about how to do it, taught me about how asyncio works (I'm deeply familiar with the model, just not in python), taught me what how Bluetooth works, what GATT characteristics are and gave me some example code that I then took and rewrote to fit my architecture.
I tried to do the normal "find a tutorial on the internet thing" first, but it's filled with worthless AI slop, ads and examples that are so trivial they're only clickbait. It's incredibly frustrating.
Meanwhile, the LLM was succinct, helpful and able to iteratively answer each question I had as learned and got deeper into it, with links to source references I could use to verify the information.
So, basically, a better search engine.
I've been writing code professionally since the 90s, and struggled through learning new things more times than I can count. This was by far the best experience I've had. It was where LLMs really shine.
But, that being said, at the end of the day I wrote the actual code myself and structured it to fit in with the rest of the architecture for the firmware.
I understood everything and when I had to pass it off to another engineer, I was able to describe everything, how it works, and why it's structured the way it is.
I think AI is even worse than drugs. On the one hand, it has some elements of a slot machine (when you have irregular success with the process, it creates a more powerful cycle of dopamine release on success). On the other hand, it makes work easier, and humans are just wired to choose the easier road. It's like walking or driving a car when you have to travel 10 miles each day. I don't see how, on a civilization scale, AI won't transform our civilization for better or worse.
Over the past six months I tried using Claude, chatgpt, Grok and Gemini. At best I got reminders of how things worked. People online say they use them to write their code. The code they supplied to me has NEVER worked or was so convoluted that I threw it away and did it myself.
At most, I use these tools as search engines. Even then some references are poor.
I'm starting to think this is becoming a sad, sad world and AI is just the new TV of the programming world.