Yes, LLMs make mistakes, but even small, self-hosted ones are pretty efficient today if you prompt them well. They’re not mind reading software so you need to be able to describe the task and HOW you want it done, not just barf in some basic instructions like “write me a copy of Facebook but better”.
Recent example: I asked AI to help me come up with a search/replace string for editing a file in vim. My prompt was something like “I am editing a text file in vim and I need to replace <description of pattern> with <description of pattern>. How can I do this with :s/ ?” And it spat out something that didn’t work. It was not valid syntax (for vim, I think it was valid standard regex, or close to it), and it didn’t quite follow the pattern I was trying to describe (although ofc that could be my fault to some extent). But it was useful in that it pointed me in the right direction to come up with a correct formula with some follow up googling and experimentation.
That has been very typical of my experiences with AI. Useful sometimes, but absolutely not “does everything to the point you don’t have to think about it” that seems to be a common opinion.
For context I’m using ~20GB local models, so not Claude, which people who pay for LLMs swear by.
Yes, older models will make those mistakes. The solution is to provide appropriate agent/skill definitions so it doesn’t just spit something out, but rather comes up with the solution, then smoke tests it in a separate scratch environment.
Claude uses this approach for reinforced learning, and it works well. Takes a few more rounds to resolve, but the solution is generally flawless (for that specific purpose).
This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
What you’re missing here is a feedback loop, validation, skills, etc. For example, if I ask one of my well-configured agents the same question, they go read the help/man page/other docs, start a vim session, quickly test/iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it, and I can keep working for the few seconds it takes. This rigor is written into my global instructions, not something you get out of the box on most models (Anthropic’s models tend to be good at this without handholding, which is part of why they’re so popular, aside from the fact that they just don’t make as many mistakes).
The same mindset scales to larger software problems, too. As long as you have a well-defined specification and good agent instructions (and/or something like Spec Kit), you can have agents break it down, implement, and then other agents compare the result to the spec, and just keep looping until it is done. The hard part is writing good specs and requirements, but that’s not a new problem.
This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
Sure, but a good web resource or even a good reference book would provide better help faster. Unfortunately it’s getting harder to harder to find those good web resources as search results get overtaken by AI slop.
iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it
I have gotten it to do this in certain situations, like the other day I wanted to simplify a math formula, so I gave it a loop with a Python script that checked its solution against the original reference version. This worked pretty well.
But I had to write code specifically for that situation. Even with a good skeleton to start with, it’s a non-negligible amount of work to get that set up. I feel like the scenario in which this is useful is kinda narrow: when I have a very good idea of exactly what I want, but some step along the way is a hassle. General software engineering, like making a whole app, is far too open ended, and most of the sub-problems I encounter in software engineering seem either too open ended or too small to benefit from this approach.
That also doesn’t account for the speed of models. My experience is a ~20GB locally hosted model takes like 1-5 minutes to produce a good length response, and the few times I have used online models they are often slower. A few minutes per iteration, accounting for debugging when it goes off track, is not exactly fast or hassle free.
I just feel like the trade off where using AI vs doing it all myself is pretty limited in when AI offers an advantage.
Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can’t screw up in, and have it not just INVENT things (“give me X”), but research the topic and base its solution on the rules created by the research.
This is what basically the Claude harness (not the local but the remote harness you can’t see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.
It sounds like you’re saying people still need to be smart enough to use them properly… Which will be a problem as people rely on them more and more, and in turn become more stupid.
I think most people are so disorganized in their thinking, they can’t “prompt” well. There’s a lot of unclarified assumptions and leaps in how many people communicate
And funnily enough, AI is great at helping streamline that process too. People just need to ASK for help (even if they’re asking the AI model) instead of being set in one way of thinking and expecting miracles.
Your experience is pretty unique then.
Yes, LLMs make mistakes, but even small, self-hosted ones are pretty efficient today if you prompt them well. They’re not mind reading software so you need to be able to describe the task and HOW you want it done, not just barf in some basic instructions like “write me a copy of Facebook but better”.
Eh idk I still have it make mistakes constantly.
Recent example: I asked AI to help me come up with a search/replace string for editing a file in vim. My prompt was something like “I am editing a text file in vim and I need to replace <description of pattern> with <description of pattern>. How can I do this with :s/ ?” And it spat out something that didn’t work. It was not valid syntax (for vim, I think it was valid standard regex, or close to it), and it didn’t quite follow the pattern I was trying to describe (although ofc that could be my fault to some extent). But it was useful in that it pointed me in the right direction to come up with a correct formula with some follow up googling and experimentation.
That has been very typical of my experiences with AI. Useful sometimes, but absolutely not “does everything to the point you don’t have to think about it” that seems to be a common opinion.
For context I’m using ~20GB local models, so not Claude, which people who pay for LLMs swear by.
I recently used it to customize a vimeopro library and I had to go through some back and forth with it but we did get there in under an hour.
I have very little coding experience, I’m a graphic designer.
I would’ve required the help of someone who’s salary would’ve been six figures to solve this problem without the Claude.
So what it’s done is make computer code much cheaper than it previously was.
I use Claude + local models.
Yes, older models will make those mistakes. The solution is to provide appropriate agent/skill definitions so it doesn’t just spit something out, but rather comes up with the solution, then smoke tests it in a separate scratch environment.
Claude uses this approach for reinforced learning, and it works well. Takes a few more rounds to resolve, but the solution is generally flawless (for that specific purpose).
This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.
What you’re missing here is a feedback loop, validation, skills, etc. For example, if I ask one of my well-configured agents the same question, they go read the help/man page/other docs, start a vim session, quickly test/iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it, and I can keep working for the few seconds it takes. This rigor is written into my global instructions, not something you get out of the box on most models (Anthropic’s models tend to be good at this without handholding, which is part of why they’re so popular, aside from the fact that they just don’t make as many mistakes).
The same mindset scales to larger software problems, too. As long as you have a well-defined specification and good agent instructions (and/or something like Spec Kit), you can have agents break it down, implement, and then other agents compare the result to the spec, and just keep looping until it is done. The hard part is writing good specs and requirements, but that’s not a new problem.
Sure, but a good web resource or even a good reference book would provide better help faster. Unfortunately it’s getting harder to harder to find those good web resources as search results get overtaken by AI slop.
I have gotten it to do this in certain situations, like the other day I wanted to simplify a math formula, so I gave it a loop with a Python script that checked its solution against the original reference version. This worked pretty well.
But I had to write code specifically for that situation. Even with a good skeleton to start with, it’s a non-negligible amount of work to get that set up. I feel like the scenario in which this is useful is kinda narrow: when I have a very good idea of exactly what I want, but some step along the way is a hassle. General software engineering, like making a whole app, is far too open ended, and most of the sub-problems I encounter in software engineering seem either too open ended or too small to benefit from this approach.
That also doesn’t account for the speed of models. My experience is a ~20GB locally hosted model takes like 1-5 minutes to produce a good length response, and the few times I have used online models they are often slower. A few minutes per iteration, accounting for debugging when it goes off track, is not exactly fast or hassle free.
I just feel like the trade off where using AI vs doing it all myself is pretty limited in when AI offers an advantage.
No amount of prompt “engineering” will help when they outright make shit up.
Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can’t screw up in, and have it not just INVENT things (“give me X”), but research the topic and base its solution on the rules created by the research.
This is what basically the Claude harness (not the local but the remote harness you can’t see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.
Like back when we had to teach people how to google shit. Tedious.
It sounds like you’re saying people still need to be smart enough to use them properly… Which will be a problem as people rely on them more and more, and in turn become more stupid.
I think most people are so disorganized in their thinking, they can’t “prompt” well. There’s a lot of unclarified assumptions and leaps in how many people communicate
And funnily enough, AI is great at helping streamline that process too. People just need to ASK for help (even if they’re asking the AI model) instead of being set in one way of thinking and expecting miracles.