A comprehensive investigation into attention map divergence, reasoning token entropy, and layer-to-layer distillation across disparate transformer architectures.
A standardized methodology and open dataset for distilling cross-architecture reasoning, code generation, and agentic traces into efficient sub-8B parameter models.