• treadful@lemmy.zip
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    I still can’t decide if these people are delusional or I am.

    Every single time I use an LLM it fucking “lies” to me or otherwise completely fails at the task. The people talking like this seem to me like they’ve never actually used it, or haven’t actually vetted the accuracy (like most AI users).

    Maybe I’m just not using the “good stuff”. Or I’m not imaginative enough to foresee a near future where these problems are actually corrected and it becomes trustworthy.

    I’ve never been so torn by a technological prediction.

    • moopet@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      15 hours ago

      Yeah, every time I say this to someone they say it’s fine when they use it because they pay for the latest expensive model. Except they’ve been saying that for the last couple of years, so I guess they were lying back then…?

    • Snapz@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      16 hours ago

      It’s a big club and you ain’t in it. They have to sell it until the end and this guy is on the outs and not getting invited to the “good” yacht parties at the moment.

      The judgement you are considering from this person is the same judgement (and morality) that thought it was fine to be a close and frequent associate of jeffrey epstein.

      Bill gates is a desperate, broken man. He is the illusion of a man you once admired vaguely and that is the only remaining currency he holds that matters to him. The best thing about bill is the wife who left him.

    • dropdrip@lemmy.ml
      link
      fedilink
      English
      arrow-up
      0
      ·
      17 hours ago

      I’m hesitant to comment because this topic can become morose. I don’t think these commentators are talking about ChatGPT or whatever dribbles out into the consumer sphere. I think their concern is the extreme dis-balance of wealth and the ability for automated systems to exacerbate that. It’s not really about capital as capital is a form of control on labour; what if you removed capital entirely as you had a more effective mechanism of control? This isn’t about computers writing your assignments or ‘taking your job’, it’s a potential magnification of what technology innately does, but to an absurd degree. Technology allows for fewer men to control more. What happens when a small group of individuals control a nation’s economy?

      ‘AI’ was used to elect Trump; ‘AI’ is manipulating stocks; ‘AI’ is being used in propaganda; ‘AI’ is forcing economic rents to increase; etc…

    • jj4211@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      22 hours ago

      Generally I’ve found that:

      If the facts are painfully obvious from a simple web search, then GenAI has a decent chance of getting it right. This can be useful if you can’t recall any “key” words well and the GenAI can craft several searches and get there.

      However, if it does mess up the facts, the result looks superficially the same as “correct”. So while it can give accurate data, you always have to double check. This can still be useful, as finding the right search terms can be a decent help.

      In coding, sometimes in some situations, you can have requirements that are absolutely testable, and thus you can have the models retry and retry until it works. This isn’t always feasible. Even when it seems feasible, you may screw up the criteria, or the GenAI when enough freedom disables a probablematic test rather than solve it, and it likely will generate code that’s not really fit to modify. There are a lot of situations where this is useful, but it is infuriating that non technical people and even some low skill technical people assume this is always the case.

      Then when you get away from facts mattering, it gets “better”. Example, someone jokingly asked for one to “make gta6”. After a while it came back with a GTA 1 clone. Lots of people were impressed, because whatever it did could be considered a success even as it obviously didn’t match the expectation. The operators also like to GenAI some webcomic, where it is a fiction. They almost always didn’t have any interesting thought going in so they tend to be crap, but “correctness” didn’t matter.

      Of course, also making fakes. Supreme case where looking correct matters but being factually correct does not matter at all. GenAI above all else “seems” correct.

    • kewjo@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      from my point of view it feels like most people are willing to trade their cognitive function for being lazy, which is a boundary i never want to cross.

      AI will produce a lot of code but most of it is pretty poor quality as most training data is going to be poor quality code, there’s just always going to be more bad code than good to begin with just due to how difficult quality code really is to produce. I’ll give a hint, good code is usually small and succinct.

      since my work started pushing AI live site issues have increased dramatically. turns out the person using AI looks like they have a ton more productivity but in reality that just shifts to whoever is reviewing the code. and to those who will say its the developer’s responsibility to review the code, yeah no shit, but if you ever work corporate you realize most don’t care as long as their managers think they are productive and just blame others for being bottlenecks.

      Overall it just enlightened me to how bad the average developer really is, but i guess obtaining mediocrity is the sacrifice to make in the name of “productivity”.

    • dwemthy@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      No matter how much I carefully structure a prompt, define specific behaviors in skills, and tweak the agent md files it will still just go do something I don’t tell it to or not do something it’s got really specific instructions for. We have to write all code by LLM now at work and I’m trying to do my due diligence to review code before putting it up for PR. 9 times out of 10 when I tell it to show me a diff before committing it silently runs git diff in the background and prints “that’s the full diff”. That’s with some basic “here’s what I want when I ask for a diff” in the base context.

      • jj4211@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        22 hours ago

        Problem is that Gates isn’t really “in the loop” and doesn’t have especially valuable insight.

        His position in tech was always a bit removed from the core technologist work, and now his exposure is a telephone game with people that are as distant from the tech as he was.

        It is really going around with no shortage of commentators spewing out guesswork, but Gates is given more credibility by virtue of his role 30 years ago.

      • treadful@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        That’s a fair point. I just don’t know if the future they see can become a reality.

        • redballooon@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          I said that a year ago, and half a year ago, too, expecting the S curve to hit and the technology to flatten out at some upper limit. But instead we got the agent loop, and models that make really good use of it, the releases only get faster and faster, and real improvements with each one, either faster and cheaper, but just as good, or actually a good deal more capable. I’m seeing signs of behavior that is more than just a good auto complete, and think there is actual intelligence.

          Even should the S curve start to turn towards slowing down now ( and it doesn’t look like it), the ceiling that it’s going towards is so high, I am with Bill Gates here. We are not prepared for what is coming.

          • treadful@lemmy.zip
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            I’m seeing signs of behavior that is more than just a good auto complete, and think there is actual intelligence.

            I suggest you be real careful not to anthropomorphize these systems. To me this sounds almost like AI psychosis.

            • protist@retrofed.com
              link
              fedilink
              English
              arrow-up
              0
              ·
              16 hours ago

              I also don’t agree that they have any actual intelligence, but this is definitely not AI psychosis. AI psychosis is a pervasive pattern of delusional thinking and not just one thought you disagree with

            • redballooon@lemmy.world
              link
              fedilink
              English
              arrow-up
              0
              ·
              1 day ago

              AI psychosis is what I accused by precious boss of, when he fabulated about replacing all his employees based on gpt-4o. What we have today is something completely different.

              Also, I’m not anthropomorphize these things. Intelligence is not a uniquely human quality.

    • MeatPilot@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      Where I am forced to encounter advanced LLMs is every company rolled out replacements for my “dumb” assistants. Like Alexa on an echo dot or Gemini on Android auto.

      What I ask them to do does not take significant processing power. I need you to set a 5min timer. Not talk to me like you’re alive. Don’t overthink things, just do simple things.

      All that extra banter they added in eats up processing power. So now it’s slower and works like shit because it’s formulating some complex response to my request to beep in 5 minutes. Just set the timer, my hands are covered in raw chicken juice.

      • aesthelete@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        You can disable Android auto’s Gemini by setting the digital assistant on your android phone to “none” under default apps.

        If you use these things more generally that might hurt you more than help, but it got rid of the bloviating idiot in the way of the old useful voice that gave me directions.

        • MeatPilot@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          0
          ·
          23 hours ago

          Thanks I’ll have to dig deeper. I figured out dumbing down Alexa but haven’t dug into Gemini yet.

          I just wish their was an alternative besides Apple and Google for car things. Sooner or later the choice of having it disabled is going to quietly go away.

    • flicker@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I’m in college and for one of my classes we actually have a prompt we feed to ChatGPT so that ChatGPT creates an excel spreadsheet for us as a step in an assignment process. The whole spreadsheet.

      The class is for healthcare and I won’t get more specific, but yes, you can use AI today to make whole ass Excel spreadsheets.

      (One of the steps in that particular assignment involves fact-checking the spreadsheet, btw, but about 95% of the time all the answers are word-for-word in the book, 4% they’re taken from the web and have more up-to-date info than the book has, and 1% is blatant lies.)

      • aesthelete@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        The class is for healthcare and I won’t get more specific, but yes, you can use AI today to make whole ass Excel spreadsheets.

        lol, it amazes me what some people are impressed by

        • jj4211@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          22 hours ago

          Yeah, so far when someone shows me a concrete example of this amazingly complicated thing that GenAI helped then realize, I think “really… that’s it?”

          It’s a decent reminder that a lot of people have more modest needs and a lot of the enthusiasm can come from that, and maybe that’s more understandable.

    • jerakor@startrek.website
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      The bar is not that AI needs to be right. It just needs to be more right than the employee willing to do the job for the price it costs.

      People have been bad at their jobs for years. Now computers can be to.

    • fonix232@fedia.io
      link
      fedilink
      arrow-up
      0
      ·
      1 day ago

      Your experience is pretty unique then.

      Yes, LLMs make mistakes, but even small, self-hosted ones are pretty efficient today if you prompt them well. They’re not mind reading software so you need to be able to describe the task and HOW you want it done, not just barf in some basic instructions like “write me a copy of Facebook but better”.

      • WolfLink@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Eh idk I still have it make mistakes constantly.

        Recent example: I asked AI to help me come up with a search/replace string for editing a file in vim. My prompt was something like “I am editing a text file in vim and I need to replace <description of pattern> with <description of pattern>. How can I do this with :s/ ?” And it spat out something that didn’t work. It was not valid syntax (for vim, I think it was valid standard regex, or close to it), and it didn’t quite follow the pattern I was trying to describe (although ofc that could be my fault to some extent). But it was useful in that it pointed me in the right direction to come up with a correct formula with some follow up googling and experimentation.

        That has been very typical of my experiences with AI. Useful sometimes, but absolutely not “does everything to the point you don’t have to think about it” that seems to be a common opinion.

        For context I’m using ~20GB local models, so not Claude, which people who pay for LLMs swear by.

        • Guy Ingonito@reddthat.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          20 hours ago

          I recently used it to customize a vimeopro library and I had to go through some back and forth with it but we did get there in under an hour.

          I have very little coding experience, I’m a graphic designer.

          I would’ve required the help of someone who’s salary would’ve been six figures to solve this problem without the Claude.

          So what it’s done is make computer code much cheaper than it previously was.

        • fonix232@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          I use Claude + local models.

          Yes, older models will make those mistakes. The solution is to provide appropriate agent/skill definitions so it doesn’t just spit something out, but rather comes up with the solution, then smoke tests it in a separate scratch environment.

          Claude uses this approach for reinforced learning, and it works well. Takes a few more rounds to resolve, but the solution is generally flawless (for that specific purpose).

        • Kabaka@lemmy.blahaj.zone
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.

          What you’re missing here is a feedback loop, validation, skills, etc. For example, if I ask one of my well-configured agents the same question, they go read the help/man page/other docs, start a vim session, quickly test/iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it, and I can keep working for the few seconds it takes. This rigor is written into my global instructions, not something you get out of the box on most models (Anthropic’s models tend to be good at this without handholding, which is part of why they’re so popular, aside from the fact that they just don’t make as many mistakes).

          The same mindset scales to larger software problems, too. As long as you have a well-defined specification and good agent instructions (and/or something like Spec Kit), you can have agents break it down, implement, and then other agents compare the result to the spec, and just keep looping until it is done. The hard part is writing good specs and requirements, but that’s not a new problem.

          • WolfLink@sh.itjust.works
            link
            fedilink
            English
            arrow-up
            0
            ·
            22 hours ago

            This is basically the same as asking someone who knows a bit about everything to recite obscure information from memory. I doubt most humans would do better.

            Sure, but a good web resource or even a good reference book would provide better help faster. Unfortunately it’s getting harder to harder to find those good web resources as search results get overtaken by AI slop.

            iterate until the result is correct, then give me a one-page document explaining what I need to know, with cited evidence — far faster than I’d do it

            I have gotten it to do this in certain situations, like the other day I wanted to simplify a math formula, so I gave it a loop with a Python script that checked its solution against the original reference version. This worked pretty well.

            But I had to write code specifically for that situation. Even with a good skeleton to start with, it’s a non-negligible amount of work to get that set up. I feel like the scenario in which this is useful is kinda narrow: when I have a very good idea of exactly what I want, but some step along the way is a hassle. General software engineering, like making a whole app, is far too open ended, and most of the sub-problems I encounter in software engineering seem either too open ended or too small to benefit from this approach.

            That also doesn’t account for the speed of models. My experience is a ~20GB locally hosted model takes like 1-5 minutes to produce a good length response, and the few times I have used online models they are often slower. A few minutes per iteration, accounting for debugging when it goes off track, is not exactly fast or hassle free.

            I just feel like the trade off where using AI vs doing it all myself is pretty limited in when AI offers an advantage.

      • treadful@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        No amount of prompt “engineering” will help when they outright make shit up.

        • fonix232@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          Yes it does. Just need to go beyond prompt. Add reinforcement loops, make it test the solution in a separate environ it can’t screw up in, and have it not just INVENT things (“give me X”), but research the topic and base its solution on the rules created by the research.

          This is what basically the Claude harness (not the local but the remote harness you can’t see) adds to the LLM what makes it so powerful and useful. Replicate those processes and even a small 4B mode will be incredibly capable.

      • Fishnoodle@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        It sounds like you’re saying people still need to be smart enough to use them properly… Which will be a problem as people rely on them more and more, and in turn become more stupid.

      • jtrek@startrek.website
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        I think most people are so disorganized in their thinking, they can’t “prompt” well. There’s a lot of unclarified assumptions and leaps in how many people communicate

        • fonix232@fedia.io
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          And funnily enough, AI is great at helping streamline that process too. People just need to ASK for help (even if they’re asking the AI model) instead of being set in one way of thinking and expecting miracles.

    • teslasaur@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I use AI all the time to parse log files for errors. Or write up simple scripts to do things that are one-offs or test of concept. Most, if not all succeed. I made a parser that translates the config of one brand of switches to another one, worked perfectly.

      In what way are you using an LLM? It sounds like you’re asking it moral questions, to which it of course can’t give you an answer in any sort of objective sense.

      • ThirdConsul@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        11 hours ago

        parse log files for errors

        Surely you don’t tell your LLM “ingest the whole 10GB log file, look for errors”?

        made a parser that translates the config of one brand of switches to another one, worked perfectly

        I believe we are all saying in the thread that minor things do work correctly. Because a config map is not a hard thing to produce.

    • Mikina@programming.dev
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I’ve been able to find a workflow where it’s mostly correct and can handle most of my gamedev related coding without making too many mistakes. I still have to actually read through the code and pay some attention to what it’s doing, and if it misses something and goes on a wild goose chase, it’s unusable and I have to start over (so someone who didn’t know what they are doing would be cooked), but whatever.

      Sure, it does require a lot of looping adversarial reviews, and my average token cost is around 3000$ a month (we have unlimited budgets and a pretty accurate tracking, and also definitely cheaper than consumer prices per token with how large company it is), which is actually more than my monthly salary, but it’s just a job, for a company and on a product I don’t really care about, and I can 1) keep slacking in my job while doing my own coding stuff and projects, and keep seeing how absolutely unreasonable the prices are if you want to get at least semi-submitable results.

      Is it worth it? Lol, no. The whole team is loosing codebase knowledge, we’re getting bottle-necked by pending PR reviews that are just stacking up and no one wants to do, so we’re not even more effective, the cost is absolutely absurd and in no way near sustainable.

      And that’s while the whole industry is in the “Uber pricing” phase, so it will get a lot worse. But yeah, if you can burn 100-200$ per a simple implementation task, then it can have a pretty usable results. And that 3000$ a month does not include our CI review bot, that does additional rounds of multi-agent council reviews.

    • Riskable@programming.dev
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      Give us an example of some of the prompts you’re using and what LLMs. I’m curious if it’s a use case difference or you’re using the dollar store’s customer service AI to try to help you with your coding homework.

      • treadful@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Here’s a very common response to anyone that suggests they had a bad time with LLMs. You just aren’t using the right model. You didn’t ask the right questions. You didn’t give enough context in your prompt.

        It’s not the fault of this infallible AI, it’s PEBKAC.

        Nonsense.

        • Not_mikey@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          I mean yeah, it’s not magic, it’s a tool and it does take some skill / know how to use them correctly, and the people making those comments could be trying to teach those skills.

          Like if someone said they had a bad time with Linux you’ll get similar questions and suggestions about their setup.

          • treadful@lemmy.zip
            link
            fedilink
            English
            arrow-up
            0
            ·
            23 hours ago

            I’ll concede that using LLMs in a useful way may take some skill. However, that’s not how these things are presented to everyone and it doesn’t reflect the reality of how the majority of people use them.

        • StabbingSky@lemmy.zip
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          Prompting an AI is very much a garbage in, garbage out type thing. Just like with any tool, you need to know how to use it properly to get the results you want.

          Hell, some of the things I’ve seen people ask AI would confuse a human too.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          Uh… I was just curious because sometimes it’s fun to see how the LLMs screw up. e.g. rocks on pizza.

          It’s not the lack of evidence presented, it’s the person that asked the question that’s the problem.

          • treadful@lemmy.zip
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            No you weren’t. You literally suggested I was using a “dollar store’s customer service AI to try to help [me] with [my] coding homework.”

            • Riskable@programming.dev
              link
              fedilink
              English
              arrow-up
              0
              ·
              1 day ago

              Using a dollar store’s AI to help with coding homework would be hilarious! Just like rocks on pizza.

              You seem to be placing me in the wrong bucket. I use AI professionally, yes, but it also pisses me off pretty regularly. I’m not some “AI Bro” or whatever TF you’re thinking. Just some guy who finds AI mistakes to be funny and I want to reproduce them.

    • realitista@lemmus.org
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I’m curious what you are using. The free versions of chatgpt have been like that for me, but even Gemini flash with extended thinking, also free for a while longer, is giving me pretty reliable results as long as there training data out there to derive an answer from. The higher (paid) Claude models will one shot most coding tasks.

      • ThirdConsul@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        11 hours ago

        I mean as long as the coding tasks are simple, and seeded with enough context, then sure. As in “create tinder clone works”, but I still have to check the output in my company codebase. We are leveraging llms a lot for code writing, full agentic pipelines, loops, all the shiny new approaches.

        The outcomes in all cases are still mediocre or below.

        We have twice as many open prod bugs as a year ago.

      • tyler@programming.dev
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        Claude can one shot tasks until you get a larger system then it completely shits itself. These models are nothing more than autocomplete, and they can’t hold large systems in their heads. Anthropic literally tried to rewrite all of bun using Claude, they said they did it and yet it still hasn’t released six months later.

        • Not_mikey@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          until you get to a larger system then it completely shits itself

          This has gotten a lot better for me by having a “send out scouts” skill that has a lower tier model agent search through the codebase before it starts to plan. Has handled my companies giant monolith pretty well and even can handle cross repo features as well.

          Yet it still hasn’t released six months later

          Claude code has been using the new rust bun for ~5 months now and has been working fine, and bun 1.4 that released in August is using the rust rewrite

          • ThirdConsul@lemmy.zip
            link
            fedilink
            English
            arrow-up
            0
            ·
            11 hours ago

            It is worth noting that the repo had multiple Anthropic employees commiting to it after the bun rewrite and before release, and that so far the only successful use cases that Anthropic or OpenAi showed are… rewrites from one language to another given sufficiently large test suites.

        • Psythik@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          By writing instructions to insist that it double verifies every (non obvious) claim with a minimum of two independent sources. I also told mine to always assume that the initial prompt is missing crucial context, and to ask as many follow-up questions as necessary until it has enough information to provide the answer to the question I’m really asking. (For speed and efficiency you can even make it give you multiple choice options to click on.) Because sometimes the problem isn’t with the LLM, but with the user asking the wrong questions.

          Using those two instructions alone, I’ve encountered considerably fewer hallucinations, and when I’m still not certain, I can simply click on the sources linked next to every single claim the AI makes.

          • ch00f@lemmy.world
            link
            fedilink
            English
            arrow-up
            0
            ·
            1 day ago

            Right, but that sounds like you implicitly trust the LLM and only verify statements when they seem off. So your intuition is the final arbiter if truth?

            • Psythik@lemmy.world
              link
              fedilink
              English
              arrow-up
              0
              ·
              6 hours ago

              No I’m saying that I don’t have to trust it because I instructed it to put two independent sources next to every claim it makes. I just check the sources.

        • Riskable@programming.dev
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          You check the links/references it gives you.

          Gemini does a pretty good job of this because it doesn’t seem to have much built-in knowledge. Instead, it just searches the Internet on your behalf and returns summarized results with links to where it got that specific information.

          I use it to search for scientific research all the time and the summaries often aren’t detailed enough so I actually click on those links. I’ve yet to encounter a situation where it fucked that up (invented links that don’t exist) but I have heard about it happening.

          So far, the summaries have seemed to be pretty spot-on when it comes to biology papers 🤷

      • treadful@lemmy.zip
        link
        fedilink
        English
        arrow-up
        0
        ·
        1 day ago

        I’ve not yet fucked with Claude. I don’t want to pay for it, and I really don’t like the surveillance aspect of these centralized systems. Mostly I’m using Gemini, whatever DDG had in their search results, and local models I’ve been fiddling with (like Qwen3.8 right now).

        All more or less garbage once I get into the details of anything on the edge of my expertise.

        • realitista@lemmus.org
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          Are you using Gemini in flash extended thinking? (Not flash light) . I haven’t had many hallucinations other than cases where the training data it needs just doesn’t exist (cases where I can’t find the answers by googling either)

        • RyanDownyJr@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          I’ve been using OpenCode with whateverthefuck free models they have listed on there and they all seem to do fine with agentic tasks like building me scripts or executables to make my work tasks easier.

          I used Gemini at the start with “Frontier Knowledge” and it seemed to do worse than the ones listed on OpenCode, but maybe that’s because i could only do like three prompts a week since I refuse to pay into an AI.

          end of the day, its just LLMs writing code for me, but I cannot see how this would be useful for a large scale code base, but also #NotAProgrammer.