• SirDimples@programming.dev
    link
    fedilink
    arrow-up
    0
    ·
    6 hours ago

    Wow, got about 9 repos of mine there and a few from my startup’s, only public ones are scraped so I say fair enough. Happy to have switched to running my own git infra a year ago

  • alcea@feddit.org
    link
    fedilink
    arrow-up
    0
    ·
    20 hours ago

    Thats why you write your projects now, download all weights locally and never look back.

    For posterity huggingface is terrible. If you need it. Even codebergs new changes are a disaster in teh making The days of “if its on the internet” are now your own responsibilty.

    Love how things are being rewritten on the fly now… Even better, lets scrape and “selfimprove”… WhatCouldPossiblyGoWrong

  • Bieren@lemmy.today
    link
    fedilink
    arrow-up
    0
    ·
    22 hours ago

    If someone wants to scrap the shitty ass code I have on GitHub, have at it. Talk about poisoning AI

  • onlinepersona@programming.dev
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    Get off of Github if you think this is a problem 🤷 There are alternatives like Forgejo (Codeberg), Gitlab, and Radicle (decentralised).

  • mindbleach@sh.itjust.works
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    Genuinely surprised it’s even opt-out.

    These companies train on Disney DVDs. Permission is not a factor. Training is transformative use, as much for counting letter frequency as for building a chatbot that can sort of code.

    • WhyJiffie@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      no, officer, you misunderstand! I’m not pirating this movie, I’m just training my intelligence on it! it is transformative use, see, I can now write you this summary!

        • Wiz@midwest.social
          link
          fedilink
          arrow-up
          0
          ·
          5 hours ago

          I would say this is a gray area of law. It hasn’t been tested yet. There are a few factors in determining “fair use”. One of those factors is commercialization, which could nullify fair use. Another is the amount you’re using.

  • Melllvar@startrek.website
    link
    fedilink
    English
    arrow-up
    0
    ·
    2 days ago

    A number of my repos are listed.

    But the weird part is that it also lists a repo I don’t recognize. The repo does actually exist on my github account, but it’s marked as “ignored”, and the description says it was automatically exported from Google Code. The code seems to be a MacOS shareware file encryption tool called “BitClamp”, published circa 2008.

    No idea how it got on my account.

      • Diurnambule@jlai.lu
        link
        fedilink
        arrow-up
        0
        ·
        1 day ago

        They got some broken Linux configuration from me, some project with many securities fails in it and a big amount of virus codes I got from the time I was hypefocusing on worms…

  • TeaWithDani@lemmy.world
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    In theory, it’s all supposed to be permissively licenced code and the opt out is more than other models give. I saw Starcoder as one of the more ethical models. I thought the underlying principals to be fair at least.

    I’m interested in this gut hostility to it regardless. Kind of shows how you can’t present LLMs in a positive angle no matter what.

    • AlteredEgo@lemmy.ml
      link
      fedilink
      arrow-up
      0
      ·
      2 days ago

      Machine learning is more than just “transformative use” and is not copying. Currently that is only like 98% true, because memorization does occur in a few cases. Like 0.8-2% and that number is probably less now 2 years later than that study. Ultimately once they fix the memorization issue and “purge” these memories and can no longer reproduce licensed code (which is mostly textbook examples and boilerplate code or very popular code) this will be true transformative learning.

      Then they do not require any more permission to read and learn from a book or from code than a human would. As long as you own a book or have the right to read something, you’re allowed to do whatever you want with the knowledge you gained.

      • TeaWithDani@lemmy.world
        link
        fedilink
        arrow-up
        0
        ·
        2 days ago

        Yeah, they aren’t supposed to scrape that stuff. That’s kind of where I’m wondering if they limited their collections.

        • Vorpal@programming.dev
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          Well, they included some MPL 2.0 repos of mine at least, but skipped others that use GPL3. Yet another that is a mirror of some otherwise lost firmware files for early 2000s wifi cards (and definitely isn’t free software) is also included.

          So possibly they filter out GPL2/3 specifically, rather than only include known permissive licenses. Which is a pretty bad way of doing it.

        • JackbyDev@programming.dev
          link
          fedilink
          English
          arrow-up
          0
          ·
          2 days ago

          Okay, I see you, I think I may have misunderstood how you were phrasing it. Thinking you were saying they were scanning all of GitHub because it’s all permissive.

          • TeaWithDani@lemmy.world
            link
            fedilink
            arrow-up
            0
            ·
            edit-2
            2 days ago

            We’ll they said they only scanned stuff that would have allowed them implicitly, now did they really respect that? Their ‘‘stack’’ is public, so people can review it. It’s opensource, people can audit it at least.

            The big giants can lift whatever they want from Github and we wouldn’t have the means to prove it. I’m sure Microsoft is using private repos as they like. It’s not a coincidence that Copilot was one of the earlier coding LLMs.

    • Viking_Hippie@lemmy.dbzer0.com
      link
      fedilink
      arrow-up
      0
      ·
      2 days ago

      Kind of shows how you can’t present LLMs harvesting peoples data without consent or even warning and making it difficult to impossible for people to avoid it in a positive angle no matter what.

      Fixed it for you.

      An LLM built from only consensually provided data would be perfectly fine, as long as it works without environmentally ruinous data centers.

      In fact, that was how EVERY LLM was to begin with, until regulatory capture and corporate impunity reached the current crescendo.

      • TeaWithDani@lemmy.world
        link
        fedilink
        arrow-up
        0
        ·
        2 days ago

        But if it’s permissively licenced, couldn’t I just copy bits and pieces for my own project without asking?

        Like I understand asking is always better and an opt in process for “the stack” or wtv would have been better received. Nevertheless, did they really have a legal obligation, rather than moral obligation, to ask given how this is licensed?

        • Axolotl@feddit.it
          link
          fedilink
          arrow-up
          0
          ·
          23 hours ago

          You can’t use GPL for LLMs if you don’t include a copy of the GPL license or link to the GPL license, you would also have to deal with conflicting licenses, somehow

        • mkwt@lemmy.world
          link
          fedilink
          arrow-up
          0
          ·
          1 day ago

          Most of those “permissive” licenses require redistributors to redistribute copies of the license texts in derivative works.

          But I bet these AI models aren’t doing that. And it’s a damn neat certainty that the vibe coders who use the AI model are not attaching a license disclosure containing every permissive licenses in GitHub. Even if their vibe coded app is arguably a derivative work.

          • TeaWithDani@lemmy.world
            link
            fedilink
            arrow-up
            0
            ·
            2 days ago

            I do not believe a permissive license has any notion of consent. You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you.

            Even in GPL, you can’t stop someone from using your code, see Gnome’s recent arguements with Mint over their usage of an old version of their Calender app.

            Others in the thread have mentioned we’ll see permissive licenses with exceptions for Ai in the near future. It’s a solution because that door is currently open.

            • AeonFelis@lemmy.world
              link
              fedilink
              arrow-up
              0
              ·
              1 day ago

              I do not believe a permissive license has any notion of consent. You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you.

              Does the LLM ever credit the original author when it spits out code?

            • Viking_Hippie@lemmy.dbzer0.com
              link
              fedilink
              arrow-up
              0
              ·
              2 days ago

              You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you

              In that case,though, you’re still respecting the wishes of the person applying that permissive license consensually, though.

              That’s VERY different from mass harvesting all data without permission (or credit) for profit, which would probably be against the terms of even the most permissive licenses.

              • TeaWithDani@lemmy.world
                link
                fedilink
                arrow-up
                0
                ·
                2 days ago

                Their claim was they followed the licenses and only used code they were implicitly allowed to use. I’d like to see if that was just astroturf and bullshit.

                Honestly, though this is one of the reasons why people prefer copyleft and avoid permissive licenses, because yeah anyone can profit monetarily off your work otherwise.

                • Viking_Hippie@lemmy.dbzer0.com
                  link
                  fedilink
                  arrow-up
                  0
                  ·
                  2 days ago

                  Their claim was they followed the licenses and only used code they were implicitly allowed to use

                  Was likely bullshit. Just like just about everything else people behind for profit LLMs say about their business practices.

                  I’d like to see if that was just astroturf and bullshit.

                  Almost certainly

  • dextro@feddit.org
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    I don’t see a problem as long as they stick to AGPL when building a product out of it

  • Impractical_Island@lemmy.world
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    I am a feudal peasant who just jumped through time with a strange man in a blue box and I don’t know what any of these terms mean. Please explain them to me as if I were an imbecile, please and thank you. No, I am not an LLM training to teach people we reconstitute in the future about the world today. I am just an ordinary shit farmer like the rest of the good people of my village.

  • iceberg314@slrpnk.net
    link
    fedilink
    arrow-up
    0
    ·
    2 days ago

    I’m not much of a programmer, but why arern’t more people just using GitLab instead of GitHub?

      • Axolotl@feddit.it
        link
        fedilink
        arrow-up
        0
        ·
        23 hours ago

        Git repos don’t really take many resources to host, look at Forgejo, also, the RAM shortage is very much artificial, there is no lack of RAM, it’s just that it’s all being sold all to a few companies

        • TeaWithDani@lemmy.world
          link
          fedilink
          arrow-up
          0
          ·
          22 hours ago

          Selfhosting Gitlab takes a bazillion RAMs lol. I think I got it up to 12gb idling without doing anything.

    • MonkeMischief@lemmy.today
      link
      fedilink
      arrow-up
      0
      ·
      2 days ago

      Still learning git but apparently there’s some “power features”, and, aside from that, it’s the same BS network effect that keeps everyone on all the abusive platforms.

      Discoverability, it’s where all the other people already stashed their code, sunk cost, etc…

      Siiiigh…

      • floquant@lemmy.dbzer0.com
        link
        fedilink
        arrow-up
        0
        ·
        2 days ago

        Self hosted gitlab is pretty nice, sucks that some features are still paywalled and it’s been getting somewhat bloated after 18.x tho