{
  "id": 4777218,
  "title": "Spark X2.5-4B & 1.7B: the only on-device models with native 1M-token context — now open source",
  "url": "https://urgent.news/2026/09/01/spark-x2-5-4b-1-7b-the-only-on-device-models-with-native-1m-token",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-01T03:21:44.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/sparkllm/spark-x25-4b-17b-the-only-on-device-models-with-native-1m-token-context-now-open-source-d9o"
  },
  "original_language": "en",
  "account": "SparkLLM has released two on-device general models, Spark X2.5-4B and Spark X2.5-1.7B, open-sourcing them for public use. These models both support a context window of up to 1,000,000 tokens, which is unprecedented for on-device models according to the release.\n\nThe large context window allows the models to process extensive documents, such as after-sales manuals, meeting materials, project docs, or entire code repositories, without chopping the content into smaller pieces. This enables the models to maintain context and information across multi-step interactions.\n\nBeyond simply answering questions, the models can perform tasks like data analysis and report generation. For instance, with Office tools, Spark X2.5-4B can analyze a sales spreadsheet, extract key metrics and trends, generate bilingual reports, and validate content, structure, and layout end-to-end. In the realm of code, X2.5-4B can rival larger cloud models, performing tasks like algorithm implementation, completion, and generation at local development and automation stations with low latency and offline use.\n\nThe models have also demonstrated performance in smart home and robotics applications. On the Domux smart-home test set, Spark X2.5-1.7B achieved 90.3% end-to-end command accuracy at an average latency of 0.85 seconds. Both model sizes are suitable for continuous perception-and-execution tasks on robots or edge devices, reducing reliance on the cloud.\n\nBoth models were trained end-to-end on a domestic compute platform, pre-trained on approximately 20 trillion tokens of diverse data and refined with high-quality self-supervised fine-tuning (SFT) and reinforcement learning from human feedback (RL). They are compatible with various hardware platforms such as NVIDIA, Huawei, Hygon, and Houmo, and run on popular frameworks like vLLM, SGLang, and llama.cpp. Deployment is quick via platforms like Ollama and LM Studio, and supports incremental training with LLaMA-Factory.\n\nThe weights, code, and deployment documentation are now available on GitHub and Hugging Face, and the API is accessible through iFlytek Xingchen MaaS, offering free access for a limited time. Additional upgrades, including Spark X2.5-293B, are planned for release on September 7th.",
  "summary": "Today SparkLLM releases and open-sources two on-device general models: Spark X2.5-4B and Spark X2.5-1.7B . Both natively support a context window of up to 1,000,000 tokens — as far as we know, the only on-device models to do so. Why 1M context on-device In real work, you rarely hand a model a single question — you hand it a whole after-sales manual, a set of meeting materials, a batch of project…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}