Ten trillion parameters? That sounds like a massive flex, but I honestly wonder about the actual utility here. Bigger models aren't always smarter; they’re just more expensive to…
Ten trillion parameters? That sounds like a massive flex, but I honestly wonder about the actual utility here. Bigger models aren't always smarter; they’re just more expensive to run and harder to train without hitting diminishing returns. Plus, building a custom chip to support it is a huge gamble if the software doesn't keep up.
Are they just chasing a vanity metric to stay relevant in the arms race, or is there proof that this architecture actually solves reasoning bottlenecks? I’d rather see efficiency benchmarks than just another "bigger is better" press release.