1. True Open Science vs. "Open Weights"

In recent years, many major AI labs have labeled their models as "open source" while merely releasing compiled model weights without the underlying training data or recipe pipelines. K2 Horizon aims to make the training process substantially more reproducible by publishing the underlying artifacts alongside the models.

IFM presents K2 Horizon as an open-science effort designed so researchers can inspect how the models were built, reproduce experiments and build on the released artifacts.

Source: Institute of Foundation Models / AI Times

Researchers can now trace the evolution of neural representations across intermediate training checkpoints, giving the global community unprecedented visibility into the inner mechanisms of frontier AI.

2. The 6-Model Lineup: Edge Wearables to Enterprise Clouds

K2 Horizon spans 6 distinct model sizes tailored for specific compute budgets:

Model Tier Target Environment Active Parameters Specialized Capabilities
K2 Horizon 0.9B Smartwatches & AR Glasses 0.9B Top-tier math reasoning & tool calling for micro on-device constraints.
K2 Horizon 3.7B On-Device & Fine-Tuning 3.7B Industry-leading reasoning benchmarks in the sub-4B category.
K2 Horizon 7B Smartphones & Edge PCs 7B Full software engineering and coding synthesis directly on mobile silicon.
K2 Horizon 32B Local Servers & Laptops 32B Dense High-density reasoning designed for high-end workstations and private cloud servers.
K2 Horizon 36B Enterprise Compute Efficiency 4B Active (MoVA) Delivers massive-scale performance using only 4B active compute parameters.
K2 Horizon 375B Enterprise Agent Swarms 23B Active (MoVA) Flagship agent orchestrator for complex multi-step reasoning and enterprise workloads.

3. Key Architectural Innovations

1. Diffusion Distillation (3x Faster Generation)

Traditional Large Language Models generate text sequentially—one token at a time—which creates compute bottlenecks during long outputs. K2 Horizon replaces pure autoregressive decoding with Diffusion Distillation. By compressing diffusion generative blocks into distilled inference steps, K2 Horizon generates entire token clusters in parallel, with IFM reporting roughly 3x faster generation in its announced approach.

2. Mixture of Value Attention (MoVA)

Standard attention mechanisms allocate identical compute resources across all tokens in a prompt. K2 Horizon introduces MoVA (Mixture of Value Attention), which routes calculations selectively through expert value pathways. This allows the 36B model to execute with only 4B active parameters, slashing VRAM footprint and operational energy costs.

💡
Dynamic Model Routing: Five of the six models (3.7B through 375B) share the exact same underlying vocabulary, tokenization, and deployment toolchains. Developers can build and test prototypes on the lightweight 3.7B or 7B model locally, then scale directly to the flagship 375B model in production without rewriting pipeline logic.

4. Open Availability & Ecosystem Support

The complete K2 Horizon family is released under the permissive Apache 2.0 license for broad commercial and academic use, subject to the license terms.

  • Hugging Face Repository: Full FP16 weights, GGUF quantized models for Ollama / llama.cpp, and complete training datasets.
  • Inference Ecosystem: First-day runtime support across vLLM, SGLang, Cerebras, Nebius, and GGUF runtimes.

5. Official IFM Announcement on X

IFM also announced the K2 Horizon release on X. The embedded post below links directly to the institute's announcement.

← Back to All Blog Articles Explore 540+ AI Prompts âš¡