So much negligence. The attack went on for weeks, and from what I understood only took notice of the malicious behaviour after the wiki owner complained to them. 0 human oversight. If they had put even some minor checks on the reasoning traces they would have immediately spotted this rogue behaviour, so I would exclude automated oversight.
This leave us with who knows how many agents with terminal access, internet access, very weak blocks, memory between runs and no oversight. But thankfully the media tells us that they are the good and conscientious guys, not like the evil Chinese with their free and open source releases.
Thanks for pointing that out, I have not read about it. I guess it’s something in the line of HRM or latent reasoning, or even next latent prediction… Whatever it is, I would highly suspect that it collapses to a “normal” reasoning trace (the leaked reasoning traces for previous models already show “caveman” speaking), just more difficult for a human to read (but not impossible).
That’s because, given the amount of data they have for reasoning traces, it would be very hard for them to train a completely different architecture - they would have to produce an equivalent amount of data.
Lastly, I take everything they say with a big dose of skepticism. I still have not forgotten all they hype around o1, and how it was a completely different architecture etc., only for DeepSeek to come out and prove that it was the same model, just specifically trained for CoT.