Dieser Bereich kann Inhalte enthalten, die nicht für alle Nutzer geeignet sind. Dazu können unter anderem Texte, Medien oder Diskussionen gehören, die als beleidigend, extremistisch, gewaltbezogen oder anderweitig belastend empfunden werden. Wenn du solche Inhalte nicht sehen möchtest, nutze bitte die jeweiligen Filter- und Meldeoptionen der Plattform oder meide entsprechende Threads/Communities.
Presuming they took all the data, a one-time deal would only be good if knowledge gathering actually stopped after the cutoff year- but for recent things like tech and news, the models have to keep learning and adding to their repositories.
The returns on that value sharply diminish, of course, but I think they’re still necessary. Which will leave everybody in a bind that is very funny.
Why do they need a deal? Can’t they just steal it like the rest of their training data?
The API restrictions and login requirements are meant to make scraping hard enough to make a deal worthwhile.
Ah that makes sense.
Probably because Reddit has lawyers, and money, and a little willingness to lock down their content. Unlike individual creators, they can actually file a lawsuit
Other than that, probably it’s a licensing agreement that makes AI trainers keep paying if they still using that dataset.