Your Transformer Can Hold Two Thoughts at Once: The New Linear Superposition Result
Run two unrelated documents through a large language model at the same time, and you assume the model keeps them
Continue readingYour Transformer Can Hold Two Thoughts at Once: The New Linear Superposition Result