• 0 Posts
  • 9 Comments
Joined 3 years ago
cake
Cake day: July 7th, 2023

help-circle

  • Absolutely there’s a difference. LLMs, when it comes to this specific task, are better. That’s why they’re being used here. It’s a job they are uniquely well suited to. They do indeed come with high hardware requirements, which is why you’re not forced to use them, and why they provide the option to offload the work to a cloud service.

    Personally I would absolutely not want to ever feed my documents into an off-device model, but the point of self-hosted software is that it does what you tell it to and they absolutely should include letting you make bad decisions.


  • I think you’re under the impression that the difference between those things is far greater than it actually is. Large Language Models work by developing statistical maps of associations. That’s discrimination. They’re a direct evolution of categorization models. The ability to associate a hash of a JPEG with “cat” is the same as the ability to associate “How are you?” with “Great, how about you?” It’s all associative mapping. LLMs are just the current leading edge of that technology. If you want to, for example, generate a list of tags that describe a document, an LLM is the best tool we currently have for doing that.

    To put it another way, what you term “discriminative AI” is “generative AI.” It’s generating a category or list of categories in response to an input. That’s not functionally different than generating a sentence in response to a sentence, it’s just an order of magnitude less complex. You can argue terminology but the technology exists on an evolutionary curve, with no real hard boundaries.


  • Paperless and Papermerge have always done a lot more than just OCR. If that’s all they were, most people can do that already on the software that comes with their scanner. The core selling point of these applications is automatic categorization, sorting and tagging, and those have always relied on machine learning tools. Literally the first thing you do after setting Paperless up is start training the AI.



  • Voroxpete@sh.itjust.workstoSelfhosted@lemmy.world[AIT] paperless-ngx 3.0.0
    link
    fedilink
    English
    arrow-up
    13
    arrow-down
    1
    ·
    11 days ago

    It reads as especially hysterical in this context, because Paperless is an automatic document categorization system, and I’m really sure what they think the automatic part of that is if it’s not “AI” of some broad description. Paperless, Papermerge et al are basically wrappers for machine learning tools and have been for as long as they’ve existed. LLMs are a natural and obvious fit for the kind of work these applications exist to do.

    This just feels like someone reading “Improved AI pathfinding” in the patch notes for a video game and screaming “OH MY GOD IS NOWHERE SAFE?!”


  • This article is seriously underselling the problem.

    Yes, all of this is accurate, and its good data for people fighting back against data centres. They produce shockingly few jobs, and the most highly paid of those will likely be experts moved in from elsewhere.

    But there’s a far bigger problem; AI data centres have no path to profitability. The cost of GPUs is so astronomically high and the depreciation so insanely fast that even if 3,000 - 4,000 data centres worth of AI compute demand materializes (it won’t; companies like Meta are already selling their own unused compute because demand is so weak), they still won’t be profitable, ever.

    The article talks about how the permanent employment offered - on paper - is in the tens to low hundreds and that’s true, but the real employment is mostly going to be zero. In five years time 99% of these projects will be abandoned before completion, or turned into skate parks.