

I wonder how the venture capital firms that fund Anthropic’s money-burning operation feel about this. “Oh look, now I know how that particular funding round’s money was actually spent. Fantastic.”
Developer and refugee from Reddit


I wonder how the venture capital firms that fund Anthropic’s money-burning operation feel about this. “Oh look, now I know how that particular funding round’s money was actually spent. Fantastic.”


The problem is that synthetic data is not fit for that purpose. The more of it you use, the worse at dealing with the edge cases LLMs get.
Think of it like this… You feed a language model a bunch of genuine human-written content. Great. Now it can produce the most likely text in a lot of cases. Word combinations that rarely appear in written language rarely get generated, so most of its synthetic data lacks those rare - but still valid - combinations.
Train it on this synthetic data, and now more outliers and rare combinations get filed off. Rinse and repeat.


It’s also unavoidable.


(Non)Human centipede.


I’m waiting for the day when desperate LLM companies start paying people to post real human content, only for those people to just ask ChatGPT to do it.
Then that’s a misunderstanding on their part. What I’m getting at is that if all of your new training data doesn’t reinforce uncommon - but factually and grammatically correct - outlier word relationships, then those outliers fade away.
There’s also an amplification issue. OpenAI has had to add a ton of instructions to their harnesses not to mention goblins because the model trained on a bunch of synthetic data when one of ChatGPT’s offered personalities would go “goblin mode.” So the more it mentioned goblins, the more that data was accidentally fed back into it, and all of a sudden no matter which personality you assigned, ChatGPT would go off about goblins.