this post was submitted on 10 Jul 2023
194 points (99.5% liked)
Technology
37747 readers
200 users here now
A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.
Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.
Subcommunities on Beehaw:
This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Seems very improbable that they scraped a pirate website with forced registration and tight daily download limits (10 books a day max?) to get content that's often mislabeled and not presented in an homogeneous way.
Probably it's just using the excerpt from Amazon (which instead with paid API access is much more easy to access) as a prompt and build on it
There's been ongoing suspicions that pirated content was used to train popular LLMs simply because popular datasets used for training LLMs do include such content. The Washington Post did an article about it.
Google's C4 dataset used for research included illegal websites. What remains to be seen is if it was cleaned up before training Bard as we know it today. OpenAI as revealed nothing on its dataset.
The sources for those websites are all being archived as a huge torrent. You don't have to download every single book one by one, if you are interested in all of them...