• 0 Posts
  • 15 Comments
Joined 2 years ago
cake
Cake day: March 21st, 2024

help-circle



  • For a FOSS project you share surprisingly low amounts of info about it: the most vague concept there could possibly be, no tech stack, no ideas, literally nothing, which only tells me that there’s no so called “initiative” that you’re a part of, it’s just you, and your attempt to make other people spend their time and resources to make the game of your dreams for you. Except, judging by the vague (probably AI-written) post here, by laggy (at least on my phone) the most basic-looking landing page you could find (which is 99% vibe-coded), you aren’t willing to put even the slightest effort to do even that properly.

    If i were ever to develop a foss game, i’d much rather avoid people like you, because what you do is less than mediocre. Judging by others’ reaction, i’m not the only one here.







  • i’ve heard of the dataset poisoning and degradation caused by llm-generated content present in the dataset myself, but i’m not sure whether it was a practical observation, or a mere experiment. And I still fail to see how new datasets are really useful for developing a new llms, or how it’s a problem for the devs to switch back to the older datasets.

    And the cornerstone stays the same: to have any significant effect on the final LLM quality, shouldn’t the poisoned (either by llm-produced content, or by intentional poisoning) data portion be… well, statistically significant?



  • except for poison to take in, it should be a pretty significant part of the dataset. Also, ngl, i’m not much informed on the topic, but aren’t all the datasets, if we’re talking about generic diffusion models and LLMs, already been formed? From what i gather, the innovation in AI mainly comes from utilizing new architectures, rather than training a model on something unique.