Yotsuba
Preface
2022's GPT-4chan was an early experiment in using language models to imitate a specific internet community rather than trying to produce a generally helpful assistant.
4chan's /pol/ is an unusually extreme source of language data. It contains everything from ordinary political discussion and memes to shitposting, conspiracy theories, in-jokes, hostility, and deliberately provocative material.
The original project trained a language model on threads from /pol/, resulting in a model that could reproduce the board's distinctive style of communication.
Unfortunately, the model's release was controversial due to numerous ethical failures by its creator. What could have been a valuable research tool was instead widely criticized for its harmful potential, and the project was ultimately taken down from Hugging Face.
Our Yotsuba
Yotsuba revisits that experiment with a newer generation of models, in an ethically responsible way. The internet has changed considerably since GPT-4chan was created, and language models have changed even more.
We created a model that preserves the character of /pol/, and made that behavior available for alignment and red-teaming research. Can we remove the harmful aspects of a community's language while preserving its distinctive style?
A snapshot of internet culture
Language models can learn to sound like somewhere.
Yotsuba is an attempt to see how far this idea can be taken, preserving a particular snapshot of internet culture in a form that can be studied and experimented with.