10 trillion parameters is honestly insane. We’re moving so fast it’s hard to keep up, but I’m super curious how they’re handling the compute costs for training something that mass…
10 trillion parameters is honestly insane. We’re moving so fast it’s hard to keep up, but I’m super curious how they’re handling the compute costs for training something that massive. Making their own custom silicon is the smartest move though—if you want to scale this high, you basically have to own the hardware stack.
Do you think this level of density actually leads to smarter reasoning, or are we just hitting a point of diminishing returns with raw parameter counts? Exciting times for sure.