Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.


  • velma@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    219
    ·
    1 day ago

    “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.

    Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing.

    Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.”

    Fuck these guys.

      • boonhet@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        7
        arrow-down
        3
        ·
        18 hours ago

        That’s the goal of most technology. To allow more things to get done in less time with fewer employees.

        CNC machine, tractor, chainsaw, printing press for some examples.

    • tangeli@piefed.social
      link
      fedilink
      English
      arrow-up
      64
      arrow-down
      1
      ·
      1 day ago

      That’s because, thus far, they get away with choosing not to distribute any of their trillions of dollars to the suppliers of the information they consume - money has only gone to the suppliers of hardware and power, and to influencing politicians and rewarding investors. That’s their choice, and they should not be allowed to continue to make that choice. Good luck to the NYT.

      • Grail@multiverse.soulism.net
        link
        fedilink
        English
        arrow-up
        18
        arrow-down
        3
        ·
        23 hours ago

        https://www.advocate.com/politics/national/new-york-times-transgender-controversy

        https://www.npr.org/2023/02/15/1157181127/nyt-letter-trans

        https://theintercept.com/2024/04/15/nyt-israel-gaza-genocide-palestine-coverage/

        NYT is transphobic and pro-genocide. I hope they go bankrupt, and papers with fewer conservative biases survive. Unfortunately, the opposite will happen. Billionaire-backed news will run at a loss to manufacture propaganda, while independent journalism dies.

      • Grandwolf319@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        8
        ·
        1 day ago

        But would it be worth it if they had to pay the creative costs for it?

        They are only worth it now cause it doesn’t include that AND it’s subsidized by investor money.

        • Rothe@piefed.social
          link
          fedilink
          English
          arrow-up
          9
          ·
          21 hours ago

          From an economic viewpoint AI companies are not worth is as it is now. OpenAI and Anthropic are hundreds of billions of dollars in debt and will never make a profit. Same goes for the AI subsidiaries of Microsoft, Meta, Google etc, except they are just leeching off of the mother companies and hiding their figures among their finances. Paying creators for their data would make very little difference in their overall finances.

          • kestrel7_7@lemmy.world
            link
            fedilink
            English
            arrow-up
            5
            ·
            16 hours ago

            This is why I’ve been arguing to anyone who will listen for like four years now. If this tech is so amazing, someone will figure out a way to make money off of it. Right? The fact that it’s been 4+ years and no one has made a damn cent off of it should be making more people skeptical. The fact that 4 years ago it was widely celebrated even though no one had a plan to make a damn cent off of it should have made more people skeptical back then. It’s weird that I have to keep arguing this with folks.

        • tangeli@piefed.social
          link
          fedilink
          English
          arrow-up
          7
          ·
          1 day ago

          I think the way they are operating now is commonly called predatory pricing.

          I don’t know if there is a potential for a profitable business that covers all the costs. Possibly the direct costs, but I think not if they had to fairly compensate those who created the intellectual property they consumed to train their models, and much less likely if they had to pay the indirect costs. They might, some day, compensate creators a little, like Google shares a pittance of their ad revenue with some creators. But it is unlikely to be anything near a living wage.

        • bad1080@piefed.social
          link
          fedilink
          English
          arrow-up
          3
          ·
          22 hours ago

          But would it be worth it if they had to pay the creative costs for it?

          no and they already admitted as much

        • BlaestEgnen@feddit.dk
          link
          fedilink
          English
          arrow-up
          4
          arrow-down
          1
          ·
          1 day ago

          AND it’s subsidised by too big a supply of compute power - Microsoft prints demand for Azure, by buying ownership stakes in openAI through the grant of Azure credits. There’s too much available compute.

    • TeaWithDani@lemmy.world
      link
      fedilink
      English
      arrow-up
      36
      ·
      1 day ago

      That’s kind of the funniest part in fact: these companies are destroying their own viable business segments. Bing was a huge growth driver for Microsoft. Less clicks is bad for them. They make more money on Bing ads than they do on LLMs. Same with Google.

      Reddit is getting crushed atm after it sold access to its data to train models. Chat bots make visiting these websites pointless, without replacing that traffic with anything they can meaningfully monetize.

      The more popular Gemini is, the less money Google will make. The market has already shown how much people are willing to pay for Ai subscriptions, and it isn’t all that much. None of these companies have found a way to make ads viable in LLMs either. They are beyond self sabotaging themselves at this point. It really is a doom loop.

    • belochka@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      1 day ago

      At least they are going to pick one between copyright and using all the Web as accumulated material for their answering machine.