• Victor@lemmy.world
    link
    fedilink
    English
    arrow-up
    18
    arrow-down
    2
    ·
    1 day ago

    We felt like he’s right about it being useful tools but the ethical aspects are hard to ignore. Yes, many ethical aspects.

    • BJW@lemmus.org
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      6
      ·
      1 day ago

      You’re likely conflating the tools with the implementation or the implementers. An open source model on local hardware has next to zero ethical concerns.

      There are MANY ethical concerns about the actions of the big players in AI, their data centers, and to some extent, the original, bootstrapped training data sourcing.

      • jj4211@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        ·
        23 hours ago

        Note a key point of contention is how the training data is used and whether it is effectively discarding copyright. If you invested time making an open source project to do something people appreciate and you get attribution as a result, you may be unhappy that a model trained on your stuff can let a user prompt up an embedded implementation of what your project does without any attribution.

        This pretty much applies to all models. No one limited training data to explicitly public domain stuff.

        • BJW@lemmus.org
          link
          fedilink
          English
          arrow-up
          1
          arrow-down
          3
          ·
          23 hours ago

          You might not appreciate it, but if it’s posted online then it’s no different from someone else learning to code from reading the project. It’s not making copies of the code, it’s just strengthening the weights on a neural network. Sure, if the code is so obscure that nothing else is like it then it’s possible to get the model to regurgitate some of it due to having so few relevant sources, but it’s very unlikely to be comprehensive enough that it’s violating any copyright. If a court finds that to be the case, for some fictitious example, then I’m certain they can find an agreeable resolution to the isolated case.

          However, none of that is justification for just writing off the technology entirely. Pandora’s Box has been opened. The genie isn’t going back in the bottle. You can’t close the barn door, all the cows already escaped. What do you think boycotting it will accomplish? What exactly is the goal by figuratively sticking your fingers in your ears and pretending the models don’t exist?

          • jj4211@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            22 hours ago

            I have seen this argument before and it doesn’t make sense even in theory.

            I used to work at a company that did open source work and also proprietary work with third party closed source code. The company didn’t let anyone who had seen proprietary code contribute to open source, because they felt once a person ‘learned’ from a proprietary codebase, then it’s too risky if similar looking code lands in a project.

            Imagine if someone saw the source code for Excel. Then sometime later they notice that Calc didn’t have a feature that Excel did, and contributed an implementation of the feature. Even if they hadn’t been looking directly at the Excel source code in the moment of implementation, you think Microsoft would be so “understanding” when they see someone that once worked on Excel contributing what could be construed as infringing?

            The AI companies also seem to acknowledge this, as they have offerings that promise not to use your proprietary code as training fodder. If it is not a risk of infringement, then why would it matter to promise that the proprietary code is kept out of training data? Though it was short lived, why would OpenAI have even made a deal to license Disney material if it’s all fair use anyway? If this sort of stuff is fair game, why do they get so pissy when other companies distill models?

            Even as the AI company’s have roughly defended this scenario, their defense should be a cause for concern for users. Generally they say that anything they do with things they can read is ‘fair use’, and when exhibits of clearly infringing outputs are given, they respond with the model only did that because the user’s prompt directed it, and thus the responsibility for infringement should be with the AI user, not the engine that produced the infringing output. So the possibility of an unwitting infringement is possible as the AI companies explicitly say it’s the fault of the user even if it happens.

            But we come to your last point, that essentially at this point, the whole thing is ‘too big to fail’ and thus the practical risk is low. Which is true. It’s just a bit disheartening that these companies are given free reign to interpret intellectual property law whichever way is convenient in the moment.

            • BJW@lemmus.org
              link
              fedilink
              English
              arrow-up
              1
              arrow-down
              3
              ·
              21 hours ago

              Thank you for taking the time to write a thoughtful, sincere response. I can tell you’ve given this some thought, and appreciate the fact you aren’t just regurgitating talking points.

              You make valid points regarding copyright law concerns, but my own perspective is that it’s an even playing field now. No one individual, group, company, etc was targeted, or unduly affected relative to any other. It could be argued it was ethically wrong to have been done at all, but since it was, and it was done to EVERYONE, then in my view it is a shared creation. Everyone is equally entitled to the resulting, from the social media users whose conversations trained the models, up through the senior engineers or CEOs of non-profit organizations whose code trained the models on syntax.

              None of it can be extracted reliably, none of it can be distilled - it is an amalgamation of information. Belonging you everyone. Saying you object to having unwillingly participated is understandable but meaningless since it cannot be undone, cannot be excised, and even if it could, the sheer amount on data means the elimination of any one source would hang an insignificant impact as to be noticable. So it’s moot. You may as well say you object to the Moon being named Luna in the past. Go for it, but it doesn’t change anything. Even if you convinced everyone it should have been called ‘Billy’ instead, you won’t change the reality of the past. Know what I mean?

              I don’t mean to be disheartening by highlighting the futility of objection, but it’s undeniable. There’s zero benefit possible, zero gain, and zero impact. Why wallow in complaints of the indelible? Instead, accept and adapt, as humans excel at.

              I’d personally like to see global tax laws that establish taxing of usage by businesses with that revenue used to fund Universal Basic Income for every person, so that the productivity and advances of this shared creation are fairly shared with everyone. I think that should be the goal of all objectors, because that is feasible, realistic and fair. Anything else is literally unrealistic. Like demanding of reality that gravity should give you special treatment. Such demands, while grand, are nonsensical.

              • tyler@programming.dev
                link
                fedilink
                English
                arrow-up
                1
                ·
                2 hours ago

                Claiming that global tax laws are realistic while banning AI isn’t is just cope. Banning AI needs nothing more than for people to realize the harm and to stop doing it. It’s exactly how we stopped using CFCs. It’s how several billion people stopped eating shark fin soup. It’s how we have begun moving away from disco fossil fuels to renewables.

      • zr0@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        10
        arrow-down
        2
        ·
        1 day ago

        Well, yes and no. Even open source models had to consume a lot of power for the training, including the use of non-owned data

        • BJW@lemmus.org
          link
          fedilink
          English
          arrow-up
          6
          arrow-down
          9
          ·
          edit-2
          1 day ago

          Many, many, many things “consume a lot of power” which is both common and benign. A whirlpool tub “consumes a lot of power” for example. So does an electric stove. And the power can be from any source, such as solar, hydro, or wind.

          That’s why I specified there’s minor issues that could be argued about the original data sources. However, that bird has flown the coop. There’s no putting that genie back in the bottle, and no way to undo it. You could argue the company’s responsible owe every single person on the planet some form of restitution, which I do, but I don’t think that qualifies as an ethical concern since it’s universal. No one person was harmed more than another, it was all publicly available data - so the public should benefit from the result. Refusing to use it only hampers yourself, at no detriment to those perceived as wrongdoers.

          • Bilb!@lemmy.ml
            link
            fedilink
            English
            arrow-up
            3
            arrow-down
            1
            ·
            22 hours ago

            I rarely see people agonize over the ethics of playing a game on high settings for hours because of the frivolous use of electricity.

            • BJW@lemmus.org
              link
              fedilink
              English
              arrow-up
              2
              arrow-down
              4
              ·
              21 hours ago

              Right? Yet using AI takes less electricity than that. The concerns are being wildly exaggerated, overblown, and I suspect it’s being done intentionally as a psychological attack on societies to hinder competing nations from faster adoption.

      • Victor@lemmy.world
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        1
        ·
        1 day ago

        You’re likely conflating the tools with the implementation or the implementers.

        I don’t think so, no?

        • BJW@lemmus.org
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          8
          ·
          1 day ago

          Okay. So, imagine I’m an indie game developer. I’ve got no artistic talent, no funds, and am just making a game that I want to play, not one I think will make any money.

          With me so far? I download a free, open source, open weights model on my laptop. I install the open source tools to allow the LLM to read my project files, and I give it a thorough description of the game I’m designing. I work on some aspect in the foreground, while in the background my GPU is utilized to achieve some task I’ve assigned the AI. When it’s done, I review the work, commit the changes, and assign it a new task.

          What are the ethical aspects you can’t ignore in my use case?

          • Victor@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            arrow-down
            1
            ·
            22 hours ago

            If you don’t intend to sell or distribute that game, the only aspects left that I can think of are:

            • you’re not helping yourself by using AI, you’re just making yourself dumber, so if your goal is simply to make a game, fine, but if you want to maintain your coding skills, it’s a bad idea.
            • the extra energy consumption, which happens regardless if you’re using your GPU or someone else’s GPU. Might even be worse if everyone uses their own GPU, because there’s a whole ass PC system built around that GPU that also consumes power. But I dunno if that’s the case, actually. I guess this is moot if you’re off the power grid with your own power source, such as full solar or something.
            • how has the open source model been trained? On what data?
            • BJW@lemmus.org
              link
              fedilink
              English
              arrow-up
              1
              arrow-down
              6
              ·
              21 hours ago

              Even if I did intend to sell the game, what change is there? Everything created was something that did not exist prior my prompt to create it.

              I’m making myself dumber in the same sense that learning to read & write weakens my memory skills, using a calculator weakens my math skills, or driving a car weakens my physical stamina. Everything is a trade off, and like in everything else, balance is the key consideration.

              Correct, I personally have a solar array with a battery system. I typically have an overage, and feed my power back into the network. Even if I didn’t, though, using my GPU for AI uses less electricity than playing a game with it, and even that is dwarfed by microwaving food, using a hair dryer, clothes dryer, etc. Power consumption is a common concern that is shared with everything, so there’s no justification to make it an ethical one just for AI.

              In my case, I’m using Qwen, and Kimi, which are both models trained by other models - essentially compressing the original training so a good analogy, if these were people, is my models were trained by professors who themselves trained themselves by reviewing the entirety of the Internet. It’s important, though, to highlight the fact that as neural networks this training is not retention of data - it’s simply adjusting the weights of simulated synapses based on the tagging of input.

              So many people seem to think that training means the model has like a compressed database of source material, but that is not the case - even on the original models that trained using pirated data, at no point is the training data incorporated into the model. It informs the model, in the same way you are informed about art by watching a Disney movie. Sure, you might be able to recreate something you saw from memory if you focused on doing that, but it’s not like you can make a perfect copy from your memory. The LLM are NOT idetic.

              • tyler@programming.dev
                link
                fedilink
                English
                arrow-up
                3
                ·
                2 hours ago

                Power consumption for AI is way higher than for games…your gpu is rarely at full usage when running a game, only really if you’re one of the few people that tries to get every single bit out of your game with the highest level graphics and even then you’re not using the full amount because the game would stutter. It doesn’t matter if an llm stutters. They use the full power of your gpu no matter what you’re doing.

                And the models being trained on other models mean that local LLMs use more energy than a model like Claude. You have all the energy of multiple data centers training one model and then you go and train a model off of that.

              • Victor@lemmy.world
                link
                fedilink
                English
                arrow-up
                5
                arrow-down
                1
                ·
                19 hours ago

                You keep making your case around a veeeeeeeeeryyyyy specific type of user. It’s hard not to think it’s actually a specific person you’re talking about lol. Hardly a common scenario. 😆

                • BJW@lemmus.org
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  arrow-down
                  3
                  ·
                  15 hours ago

                  Yes, it’s me! I’m the very specific type of user, and I’m being swept up in the discrimination in the mad rush to hate everything AI.

                  • tomalley8342@lemmy.world
                    link
                    fedilink
                    English
                    arrow-up
                    7
                    arrow-down
                    1
                    ·
                    14 hours ago

                    You say you are being swept up like some sort of unrelated bystander but it doesn’t look like your justification is too different from the typical ai user’s justification.