If LLMs write most of the code, the interesting compiler work moves from codegen to verification
Language design was always about human ergonomics. What changes when the main consumer is a model?
I'm one of those people who likes to hand craft code and put a lot of attention and intention into style, so I say this without any enthusiasm: I don't think writing software post-LLM is going to be the same as pre-LLM. Writing code by hand stays important, but I think it becomes niche, the way reading and writing assembly is important but niche.
What actually interests me is the second order effect on language design.
Language features were always chosen for the developer. Code isn't only instructions for a machine, it's in some way a specification for other developers and for your future self. So the metric was developer experience, which is subjective, which is why we've been arguing about languages for fifty years without settling anything.
If the main consumer becomes a model in a loop, that metric stops being subjective. You can measure it.
My intuition is that this is very good for functional languages, and not for the usual reasons. The usual argument is that pure code is cleaner. The argument I care about is that you generate 10K lines, you read 100 of them somewhere in the middle, and you can be sure nothing else is affected. That property is worth almost nothing when a human reads 100 lines a day. It's worth a lot when the bottleneck is how much you have to load into a context window before you can safely change anything. Purity is basically context compression.
Which makes me think the interesting question isn't the language, it's the compiler. Two things I can't stop thinking about.
First, the compiler's most valuable output is now the error message, not the binary. In an agent loop the thing that decides whether it converges is the quality of the feedback signal, and every diagnostic system I know of was designed for a person reading a terminal. Rust's errors are beautiful for a human: spans, underlines, colors. For a model they're expensive and lossy. One bad annotation produces forty cascading errors and you pay context for all forty to learn one bit of information. I don't see anyone designing diagnostics for a machine consumer: dedup, root cause first, and ideally a counterexample instead of a complaint. "Expected Foo, found Bar" tells you much less than "here's an input where this function breaks."
Second, the query I keep wanting and have never seen exposed anywhere: blast radius. If I change this function, what is the set of things whose behavior can change? That's a compiler question. It's answerable with a decent module system and some notion of effects, and it's basically hopeless in a language where anything can mutate anything. Batch compilation is the wrong shape for it anyway, because the agent doesn't want an artifact, it wants an answer to a question. rust-analyzer and Salsa look closer to the right architecture than a normal compiler driver does.
Counterargument I don't have a good answer to: models are much better at Python than at OCaml or Haskell and it isn't close, because the corpus is what it is. So the languages with the best properties have the worst priors. Maybe RL against a compiler fixes that, since a sound checker is close to a free correctness reward and you can't get that in Python. But I haven't found real numbers either way.
The other thing, and this is the part I find crazy, is the trend of people not reading the code at all. I don't have a principled argument against it yet, other than that if the code isn't the artifact you review then something else has to be, and I don't see what that something is.
Anyway. Does blast radius exist as a first class query in any toolchain? And is anyone designing compiler output for machines, rather than serializing human error messages to JSON? #technology source