Coding AI without deterministic outcomes
I recently had a conversation with an engineer on our team about working with AI. He was frustrated with some of the work he had been doing on AI features, and a lot of what he was saying resonated with me. He was putting into words things that I had felt, and that other people had described to me as well.
Using AI to help write code has already altered some of the ways engineers have traditionally found satisfaction in their work. People took pride in the ability to write code really quickly, create clever algorithms, or elegantly organize their code. AI does a lot of that work now.
But there is another source of satisfaction in software engineering: completion. I made it, it works, I gave it to other people. Every time I finish a feature — be it a tiny one that took an hour or a big one that took a month — when I get to the end, I feel satisfaction.
Using AI to do that work can actually be more satisfying. The rate at which I get the bigger satisfaction of achieving more work actually got better.
Coding AI is different. By that I mean building features that use an LLM, where much of the programming involves writing prompts to get the output you want. For me, the frustrations are a balance of both these things: I do not know exactly how to make it do what I want, and I do not know when I am done.
You cannot trace the output back to the input
Let me start with the first frustration: I do not know how to make the LLM do what I want.
A lot of the work involved in building an LLM feature is writing the prompts that will get it to do what you want. But we do not talk to the LLM with code. We talk to it with natural language, in the same way that you prompt Elle (the AI assistant in Aha! software), Claude, or ChatGPT.
English is not a deterministic way of communicating. You write in a way that you think might cause the LLM to do what you want, and then you run it. Let's say it kind of does what you want, but not exactly. I now have a problem. How do I fix what I told it to do so it does exactly what I want? I do not know. Do I need to be more persuasive? Should I be argumentative? Do I need to be rude? Was I ambiguous in a way that I did not understand as I wrote it?
There are all these different ways that I might have gotten it slightly wrong. But most importantly, it is not deterministic to figure out how to fix it.
This contrasts with writing normal software engineering code. When it does not work, I can see what happened. I can go back and look at my code and say: For that to have happened, logically it must have done this. I then look at the code and see where the logical flaw was. Or I use a debugger that gives me even better visibility into exactly what happened.
With an LLM, I change the language a little bit and something different happens. It is still not quite what I want, so I change the language again and something different happens. I cannot look at the output of the LLM and say, "It did that exactly because I did this."
Nobody knows. That is the mystery of it. We cannot look inside the model and trace an output back to a specific instruction or decision. Even the people running the model cannot fully do that.
In that respect, working with an LLM is much more like working with another person than working with traditional software. You say things, and they do things and say things back. Sometimes what they do and say back is exactly what you wanted or expected, and other times it is not. You cannot look inside their head to figure out why. They might not even know themselves why they reacted the way they did. It is unknowable.
'Good enough' replaces done
The other frustration is that I do not know when I am done. With deterministic programming, you can usually specify ahead of time what correctness is. Maybe I expect a light to turn red. I can write code, test it, iterate, and debug until I make it do exactly what I expected it to do. If the light is red, I am done. It cannot be better. There is no reason to continue.
With an LLM, the output is often natural language. What does it mean for that language to be correct? There are many variations that would achieve the purpose. There is no perfect answer.
The problem is a little bit more insidious: I do not know when to stop. I am never done.
An LLM answer might be good enough. But could it be better? Perhaps. At some point, I have to say, "Good enough." And good enough is not the standard we normally hold ourselves to.
Testing does not necessarily solve this. If the output is text, it is really difficult to write a test in code because English is flexible. In fact, the way to test LLM outputs is to use another LLM. But now, all I have done is layer the same problem a second time. How do I know if my test is good enough?
It is turtles all the way down. As an engineer, it did not become any more satisfying. I just have the same problem a second time.
Prompting is a skill. Like any skill, there is practice and learning, and you can definitely get better at it. The way I think about it when I am working with an LLM is this: If I were trying to persuade a person to do this, what would I tell them? I would not talk at them for an hour. They are going to forget. I need to be crisp and brief.
But even building up that skill does not solve the fundamental problem. There is still no determinism to be found. So I have been thinking to myself: What should we do about this dilemma?
A better fit for some engineers
I have not come up with a technical solution to this. I do not think there is a way to make LLM work deterministic.
But there are people who get their greatest joy from interacting with other people. I have seen engineers move into engineering management because that is where they were most energized. Many technical product managers started as engineers too, but found that they enjoyed thinking about the product holistically and working with a team to bring it to fruition.
Perhaps AI feature work will prove to be a better fit for some of those people. We already see engineers at Aha! who genuinely enjoy figuring out how to get an LLM to do what they want.
Like every part of engineering, people have different preferences. Some engineers love building user experiences. Some would much rather deal with database or backend work. We try to steer people toward the types of projects that bring them the most joy, because that is where they are most productive and happiest.
Programming AI features is another category like that. For some engineers, figuring out how to persuade an LLM will be a craft of its own.
Prefer AI to do the building for you? Get a free trial of Aha! Builder — the enterprise AI app building tool.
