undefined | Better HN

0 pointsdjeastm11h ago0 comments

When it was Copilot tab-completing lines, people would say, "yea, but you still have to make sure you're the one writing the whole functions".

Then when it was completing functions, people would say, "yeah, but you still have to make sure you're the one writing the logic around the functions"

Then when it was completing the logic around the functions, people would say, "yeah, but you still have to make sure you're the one writing the features"

Now it's completing features and people say, "yeah, but you still have to make sure you're the one writing the architecture"

I don't know if architecture is a solvable problem for these models, but it is interesting watching the expectations moving over time.

0 comments

raincole11h ago

The "people" in your hypothetical story have been wrong the whole time. The correct attitude is:

When AI can complete lines, you still have to read and understand the code.

When AI can complete whole functions, you still have to read and understand the code.

When AI can complete features and tickets, you still have to read and understand the code.

brightball9h ago

I heard a talk from a VP at NVIDIA a couple of months ago and he echoed this. Essentially their policy is "you are still fully responsible for the code you ship, whether AI helps with it or not"

rwmj9h ago

While also no doubt telling C-level that the AI can write code completely automatically.

The developers in this scenario are there to absorb the blame when things go wrong. "Human crumple-zones" to protect the company.

aeciorc4h ago

this is a good policy, as long as the productivity expectations match it. The problem happens when you combine "you're responsible for what you ship" with "you need to be 100x faster"

cynicalpeace8h ago

Culpability is and always will be the limiting factor in AI adoption.

As humans, we need another human to blame when things go wrong.

Especially in situations that are catastrophic when things go wrong.

9dev8h ago

Does anyone out there working with AI coding agents not have this policy?

3 more replies

senordevnyc9h ago

I think being responsible for the code is a better framing. I run a saas and I don’t always review all the code, but this thing supports my family, so I am acutely aware that I’m responsible for what it does. My customers aren’t going to let me blame the agent for fucking up their workflows.

But that still doesn’t mean I review all the code. I tend to review defensively, based on the potential for harm if this piece of code is broken. And I rely a lot on tests, static analysis, canaries, analytics, health checks, etc. to reduce risk for when I’m wrong. So far it’s working.

ventana5h ago

> you still have to read and understand the code

Which is a very similar approach to any serious code. If you just hired a very clever, enormously knowledgeable intern, and they wrote a bunch of code for you overnight, you would probably review it.

Yes, in some cases, either hobby projects or throwaway code, you could just take it and use it as is, and I surely do, for the code no one cares about. But at work, I would rather review it.

globnomulous7h ago

Precisely. And this is why all the MCP servers that people at my company are writing aren't worth using: their apparent goal is automate as much as possible. They're encouraging people not to pay attention. This results in bad code, bad tests, and bugs.

jstummbillig11h ago

Not at all. Code is not important, intent is. The leader of a product/company does not have to read code. It doesn't matter if it is generated by humans or non-humans. It simply needs to be correct enough to be usable and then steerable towards better outcomes. Understanding of code never existed from the business perspective.

throwaway17373811h ago

It does in safety critical industries. You can get grilled by regulators about your source code. And lawyers will use it as evidence in court.

roncesvalles5h ago

The code codifies the intent and is the long-term source of truth for how your business actually operates.

>The leader of a product/company does not have to read code.

That's because he's paid a bunch of people 300k to read it and make sure it aligns with the company's objectives and interests. Part of the reason why devs are paid so much is because they're literal business administrators for some narrow slice of the company's operations. The devs are the leaders that you're referring to.

Even in multi-hundred-billion-dollar companies there are so many mission critical things that are owned by just 2 SWEs.

raincole10h ago

> The leader of a product/company does not have to read code.

Yeah, because they believe (sometimes wrongly) their subordinates read it.

> Understanding of code never existed from the business perspective.

It does, it's called organizational wisdom and domain knowledge, because you need those witty names to sell books to aspiring managers.

contagiousflow8h ago

Can you think of a good way to encode intent into a system?

herdcall11h ago

I'm no longer sure you have to, actually. I mean, we do trust the assembly that compilers produce without having to read it, don't we? We're rapidly getting to that stage with LLMs, IMO.

bigfishrunning11h ago

The assembly is a deterministic transform of the input logic, and if it doesn't match then it's a bug in the compiler. If an LLM-based code generator doesn't match what you asked for, that's OK, just pull the slot-machine handle again. that's the difference.

4 more replies

gregsadetsky10h ago

I know it’s tiring to talk about “hallucination”, but truly, models still do hallucinate

They constantly say they did a thing they didn’t, say they know how to solve something when they don’t, etc. Regardless of guard rails or tests - AI forces a constant vigilance of a new kind.

Not just “what might have gone wrong” but also “what do I think is working but isn’t actually”.

And we’re not even talking about how it chooses substandard solutions, is happy to muddy code/architectures, add spaghetti on top of spaghetti etc.

Agentic coding often feels like an army of unexperienced developers who are also incredibly eager to please.

Zardoz8410h ago

"still" isn ghecorrect word. They always be having hallucinations

1 more reply

amw-zero10h ago

This is a really, really, really bad comparison. I used to say the same thing. But the semantic distance between compiling a for loop to equivalent assembly instructions is much smaller than the distance between "I'd like a web application that can store and retrieve todo items." The space of the latter is practically infinite in what can be "compiled."

1 more reply

SpaceNoodled10h ago

I've actually taken to double-checking the assembly in some instances. There are surprising times that the compiler won't make the shortcuts and optimizations you thought it should, and I also used this method to call out an unsuitable compiler since I caught it spitting out some ridiculous 10x-long set of instructions in certain critical instances.

bayindirh10h ago

> we do trust the assembly that compilers produce without having to read it

Yes, because wrong assembly blows really loudly. From wrong behavior to invalid instruction errors and everything between them. Moreover, compilers are battle tested over the years, with extremely detailed test suites, and extreme testing (everyday, hundreds of thousands users test and verify them).

Also, as people said, assembly generation is deterministic. For a given source file and set of flags, you get the same thing out. Byte by byte, bit by bit. This is what we call "reproducible builds".

AI is not like that. It's randomized on purpose, it pulls from training set which contains imperfect, non-ideal code. "Yeah, it works whatever", doesn't cut it when you pull a whole function out of its connections, formed by the training data. It can and will make errors, because it's randomized from a non-ideal pool.

Next, sometimes you need tight code. Fitting into caches, running at absolute performance limit of the processor or system you have. AI is not a good fit here. Sometimes you go so far that you optimize for the architecture at hand, and it works slower on newer systems, so you need to re-optimize that thing.

For anyone who reads and murmurs "but AI can optimize", yes, by calling specific optimization routines written by real talented people for some cases; by removing their name, licenses, and context around them. This is called plagiarism in its mildest form and will get you in hot water in academia, for example. Writing closed source software doesn't make you immune from cheating and doing unethical things.

Lastly, this still rings in my ears, and I understood it over and over as I worked with more high performance, correctness critical code:

I was taking an exam, there's this tracing question. I raise my head and ask my professor: "Why do I need to trace this? Compiler is made to do this for me". The answer was simple yet deep: "If you can't trace that code, the compiler can't trace it either".

As I said, I just said "huh" at the time, but the saying came back and when I understood it fully, it was like being shocked by a Tesla coil.

Get your sleep, eat your veggies and understand your code. That's the four essential things you need to do.

eska7h ago

We don’t. That’s why tools like godbolt are popular, debuggers can jump into assembly, and compilers can output assembly files.

yakattak10h ago

I want to preface this with that I am all for agentic engineering.

I am so tired of hearing about this false equivalency. Compilers are deterministic, their outputs are well understood and they’re transparent.

LLMs are not.

1 more reply

aprilthird202111h ago

We are not rapidly getting to that stage with LLMs and frankly it's hilarious that you are claiming so.

For anything other than Greenfield, new code projcets without dependencies and conventions and connections to other proprietary code, it has to be reviewed. Even for that case it's not good to not review code

bluGill11h ago

The models can do architecture. However they typically (at least currently) do a really bad job until you force them. I use AI all the time, it is getting better, but I still review every single line. Individual lines are no today are not better than tab completion of last year - sometimes really good and save my typing, sometimes really really bad.

embedding-shape11h ago

Anyone who understands the motivation, reasoning and goals can do the architecture. The crux is that hardly anyone actually understand those and even less is aligned on those, that's when misalignment happen over time, LLMs or not.

Considering how fast we can poop out code now, I think this issue is just more visible than before, but it's been an issue for as long as I've been a developer. Almost no one knows what they actually want, and half the job is trying to coax out what they want to be able to do, so you can properly architect it.

the__alchemist10h ago

> I don't know if architecture is a solvable problem for these models, but it is interesting watching the expectations moving over time.

I think the solution is between the lines of this article. The author states the steps leading to this, but doesn't arrive at it explicitly. It has been obvious (With 50/50 hindsight) to me since LLMs started getting popular, and holds:

LLMs are fantastic for software dev. If you don't let it write architecture. Create the modules, structs, and enums yourself. Add as many of the struct fields and enum variants as possible. Add doc comments to each struct, enum, field, and module. Point the LLM to the modules and data structures, and have it complete the function bodies etc as required.

winwang5h ago

Yeah, I pretty much agree. Opus and GPT will both come up with the most "organically-grown" "designs" if you let them. They do slightly better when asked to design first, but they seem to avoid many important questions (and definitely skip asking the user much of anything at all). I can only say it feels they "want" to ship as fast as possible while assuming I'm not going to actually review the PR.

onlyrealcuzzo7h ago

> I don't know if architecture is a solvable problem for these models, but it is interesting watching the expectations moving over time.

At least with current languages, I think the primary problem is they are globally complex, and it's not scalable for them (and certainly for you to review a codebase they've mainly or completely generated) that the invariants you want are being withheld.

No matter how many times you tell them - there is ZERO blocking allowed on the critical path, they will add blocking on the critical path.

No matter how many times you tell them any time they do X, they need Y type of test, they will do X without Y type of test.

They cannot follow directions 100%. Neither can people.

But they are more random. The mistakes people make are less likely to do the exact polar opposite of what you wanted to do.

People are less likely to see a critical invariant in the code, build themselves a loophole to get through it, write a test that the code fails successfully, and then tell you they did exactly what you asked for, and burry it in a 5k line commit, where 1000 lines are them changing comments that shouldn't be there in the first place.

LLMs are great. I'm convinced they're the future. I'm building a language specifically for them: https://GitHub.com/Cuzzo/clear - and to make it easier for YOU to work with them.

I think once we get around this language problem, that they need global context for things where they shouldn't, it will be a challenge to work with them.

I've had success with them, but it's been so frustrating, that I question how much it's been worth my sanity.

mdswanson5h ago

I refer to this as "disposable architecture." Not that architecture doesn't matter, but that the architecture that worked yesterday doesn't necessarily need to be the architecture that works today.

jayd167h ago

Are any of these steps actually solved? AI tab completion still kinda sucks.

They can keep internal consistency so the more you let it write the more it can write with internal consistency. It still fails at all of these levels as soon as you are looking at each level of detail.

indoordin0saur5h ago

It's even farther along than you think. It's the one writing the comments you're responding to. So why are you still thinking up and typing out your HN comments?

wolttam10h ago

These models understand architecture perfectly well, but they're not trained to care about it when being asked to complex X or Y feature. They're trained to implement the feature in the shortest route possible.

So it's not much of a surprise that this is the situation folks find themselves in with the current models.

wiseowise8h ago

> but it is interesting watching the expectations moving over time.

While the salary stays stagnant or even reduced if you adjust for inflation.

vrganj11h ago

As somebody with a colleague that is using AI agents to "complete features", let me tell you, it is not. It is taking that dude so much longer to prompt and reprompt and then prompt again until it is anywhere close to something that passes review than it would take any competent mid-level engineer to just build the whole thing with some autocomplete help.

Have people's standards for quality just completely vanished in the pursuit of the shiny new thing? Is that guy doing something wrong?

That has also been my experience with this sort of thing fwiw, which is why I gave up and do more of a class-by-class pairing with an LLM as a workable middle ground.

snowe20107h ago

weirdly made up scenario. I'm the person in the very first sentence. Tab-completing lines is still dog-shit. The majority of the time it has no clue what I'm going to write. Just because it can now write a lot more stuff doesn't mean it isn't still just as incorrect.

Also, you've set up a huge strawman here. Who are these people saying these things in this order and why is that the argument and not "You need to be reviewing every line of code that gets written and understand it."

Your argument is nonsense.

koumou9210h ago

100% agree. Obviously AI is at a point where the developer has to do the architecture. Or at least be in control of what kinda of architecture the AI is implementing. You can't one-shot huge features in huge codebases with AI. You are bound to get strange decisions. But that does not mean they are not worth using. That's a silly take.

hansmayer11h ago

> Now it's completing features

It's completing shit. Even if it does not implement some lazy stuff with empty catch blocks (i.e. happy path from programming 101 tutorials), it will either expose your secrets in a sensible place or do some other stupidity.

user3428311h ago

I felt the same with:

"it takes too much effort to get the output production ready"

turning into

"maybe long term the maintenance will be more expensive"

I give it three months until people realize that you rarely need to review every single line and fully understand the code, like so many comments are claiming.

camdenreslink11h ago

If you work on a product that has an existing user base that has an expectation that things will still work then you definitely still need to read the code. LLMs frequently break things or introduce subtle incompatibilities.

Maybe on projects with no users you can yolo things.

user342837h ago

It's not about the number of users but the kind of software you develop.

In a mobile app, do you think it's more important to test that your drag gesture works as expected on the phone, or to understand every line of the implementation?

keybored10h ago

There are always people who will disagree, no matter how amazing something is, and they naturally respond with concerns close to the locus of the LLMification. It would be absurd to respond to “AI autocomplete is great now” with “but you still need to architecture your code”. What’s people saving seconds writing code minutiae got to do with architecturing the code?

This blob of people criticizing AI is just that, a blob. A gaggle of discrete people that your brain makes up a narrative about being some goalpost shifting entity.

Of course there could be individuals who have moved the goalposts. Which would need a pointed critique to address, not an offhand “people are saying” remark.

dzonga11h ago

the autocomplete can be shit some times.

callamdelaney9h ago

Architecture is one of the easiest things in programming frankly.

taytus10h ago

Nice Fiction story.

j / k navigate · click thread line to collapse

0 comments

raincole11h ago

The "people" in your hypothetical story have been wrong the whole time. The correct attitude is:

When AI can complete lines, you still have to read and understand the code.

When AI can complete whole functions, you still have to read and understand the code.

When AI can complete features and tickets, you still have to read and understand the code.

brightball9h ago

I heard a talk from a VP at NVIDIA a couple of months ago and he echoed this. Essentially their policy is "you are still fully responsible for the code you ship, whether AI helps with it or not"

rwmj9h ago

While also no doubt telling C-level that the AI can write code completely automatically.

The developers in this scenario are there to absorb the blame when things go wrong. "Human crumple-zones" to protect the company.

aeciorc4h ago

this is a good policy, as long as the productivity expectations match it. The problem happens when you combine "you're responsible for what you ship" with "you need to be 100x faster"

cynicalpeace8h ago

Culpability is and always will be the limiting factor in AI adoption.

As humans, we need another human to blame when things go wrong.

Especially in situations that are catastrophic when things go wrong.

9dev8h ago

Does anyone out there working with AI coding agents not have this policy?

3 more replies

senordevnyc9h ago

ventana5h ago

> you still have to read and understand the code

Which is a very similar approach to any serious code. If you just hired a very clever, enormously knowledgeable intern, and they wrote a bunch of code for you overnight, you would probably review it.

Yes, in some cases, either hobby projects or throwaway code, you could just take it and use it as is, and I surely do, for the code no one cares about. But at work, I would rather review it.

globnomulous7h ago

jstummbillig11h ago

throwaway17373811h ago

It does in safety critical industries. You can get grilled by regulators about your source code. And lawyers will use it as evidence in court.

roncesvalles5h ago

The code codifies the intent and is the long-term source of truth for how your business actually operates.

>The leader of a product/company does not have to read code.

Even in multi-hundred-billion-dollar companies there are so many mission critical things that are owned by just 2 SWEs.

raincole10h ago

> The leader of a product/company does not have to read code.

Yeah, because they believe (sometimes wrongly) their subordinates read it.

> Understanding of code never existed from the business perspective.

It does, it's called organizational wisdom and domain knowledge, because you need those witty names to sell books to aspiring managers.

contagiousflow8h ago

Can you think of a good way to encode intent into a system?

herdcall11h ago

I'm no longer sure you have to, actually. I mean, we do trust the assembly that compilers produce without having to read it, don't we? We're rapidly getting to that stage with LLMs, IMO.

bigfishrunning11h ago

4 more replies

gregsadetsky10h ago

I know it’s tiring to talk about “hallucination”, but truly, models still do hallucinate

They constantly say they did a thing they didn’t, say they know how to solve something when they don’t, etc. Regardless of guard rails or tests - AI forces a constant vigilance of a new kind.

Not just “what might have gone wrong” but also “what do I think is working but isn’t actually”.

And we’re not even talking about how it chooses substandard solutions, is happy to muddy code/architectures, add spaghetti on top of spaghetti etc.

Agentic coding often feels like an army of unexperienced developers who are also incredibly eager to please.

Zardoz8410h ago

"still" isn ghecorrect word. They always be having hallucinations

1 more reply

amw-zero10h ago

1 more reply

SpaceNoodled10h ago

bayindirh10h ago

> we do trust the assembly that compilers produce without having to read it

Also, as people said, assembly generation is deterministic. For a given source file and set of flags, you get the same thing out. Byte by byte, bit by bit. This is what we call "reproducible builds".

Lastly, this still rings in my ears, and I understood it over and over as I worked with more high performance, correctness critical code:

As I said, I just said "huh" at the time, but the saying came back and when I understood it fully, it was like being shocked by a Tesla coil.

Get your sleep, eat your veggies and understand your code. That's the four essential things you need to do.

eska7h ago

We don’t. That’s why tools like godbolt are popular, debuggers can jump into assembly, and compilers can output assembly files.

yakattak10h ago

I want to preface this with that I am all for agentic engineering.

I am so tired of hearing about this false equivalency. Compilers are deterministic, their outputs are well understood and they’re transparent.

LLMs are not.

1 more reply

aprilthird202111h ago

We are not rapidly getting to that stage with LLMs and frankly it's hilarious that you are claiming so.

bluGill11h ago

embedding-shape11h ago

the__alchemist10h ago

> I don't know if architecture is a solvable problem for these models, but it is interesting watching the expectations moving over time.

winwang5h ago

onlyrealcuzzo7h ago

> I don't know if architecture is a solvable problem for these models, but it is interesting watching the expectations moving over time.

No matter how many times you tell them - there is ZERO blocking allowed on the critical path, they will add blocking on the critical path.

No matter how many times you tell them any time they do X, they need Y type of test, they will do X without Y type of test.

They cannot follow directions 100%. Neither can people.

But they are more random. The mistakes people make are less likely to do the exact polar opposite of what you wanted to do.

LLMs are great. I'm convinced they're the future. I'm building a language specifically for them: https://GitHub.com/Cuzzo/clear - and to make it easier for YOU to work with them.

I think once we get around this language problem, that they need global context for things where they shouldn't, it will be a challenge to work with them.

I've had success with them, but it's been so frustrating, that I question how much it's been worth my sanity.

mdswanson5h ago

I refer to this as "disposable architecture." Not that architecture doesn't matter, but that the architecture that worked yesterday doesn't necessarily need to be the architecture that works today.

jayd167h ago

Are any of these steps actually solved? AI tab completion still kinda sucks.

indoordin0saur5h ago

It's even farther along than you think. It's the one writing the comments you're responding to. So why are you still thinking up and typing out your HN comments?

wolttam10h ago

So it's not much of a surprise that this is the situation folks find themselves in with the current models.

wiseowise8h ago

> but it is interesting watching the expectations moving over time.

While the salary stays stagnant or even reduced if you adjust for inflation.

vrganj11h ago

Have people's standards for quality just completely vanished in the pursuit of the shiny new thing? Is that guy doing something wrong?

That has also been my experience with this sort of thing fwiw, which is why I gave up and do more of a class-by-class pairing with an LLM as a workable middle ground.

snowe20107h ago

Your argument is nonsense.

koumou9210h ago

hansmayer11h ago

> Now it's completing features

user3428311h ago

I felt the same with:

"it takes too much effort to get the output production ready"

turning into

"maybe long term the maintenance will be more expensive"

I give it three months until people realize that you rarely need to review every single line and fully understand the code, like so many comments are claiming.

camdenreslink11h ago

Maybe on projects with no users you can yolo things.

user342837h ago

It's not about the number of users but the kind of software you develop.

In a mobile app, do you think it's more important to test that your drag gesture works as expected on the phone, or to understand every line of the implementation?

keybored10h ago

This blob of people criticizing AI is just that, a blob. A gaggle of discrete people that your brain makes up a narrative about being some goalpost shifting entity.

Of course there could be individuals who have moved the goalposts. Which would need a pointed critique to address, not an offhand “people are saying” remark.

dzonga11h ago

the autocomplete can be shit some times.

callamdelaney9h ago

Architecture is one of the easiest things in programming frankly.

taytus10h ago

Nice Fiction story.

j / k navigate · click thread line to collapse