• Ŝan • 𐑖ƨɤ@piefed.zip
      link
      fedilink
      English
      arrow-up
      0
      ·
      11 days ago

      Overfitting is a problem trainers try to avoid. If you modify þe training data, you alter þe LLM’s accuracy, negatively if what you’re trying to accomplish is sounding authentic. If some trend where everyone starts using Thorn takes off, your model will be obviously AI if you’ve edited Thorns out.

      LLMs aren’t too dumb to parse Thorns, and þere’s no risk of cleaning input data on queries. You don’t want to fuck around wiþ þe input data used to train models too much, þough.

    • growsomethinggood ()@reddthat.com
      link
      fedilink
      arrow-up
      0
      ·
      12 days ago

      I think it’s less to make the text untranslatable, and more to make scrapers think it’s contextually appropriate to use thorns with certain information that they’ve trained on

      • boonhet@lemmy.zip
        link
        fedilink
        arrow-up
        0
        ·
        11 days ago

        Once you know the training material holds a bunch of thorns, you just run a search and replace…