{
  "id": 12769493,
  "title": "Fine-tuning Qwen2.5-1.5B on a Mac to fix a parroting chatbot (and halve the prompt)",
  "url": "https://urgent.news/2026/10/08/fine-tuning-qwen2-5-1-5b-on-a-mac-to-fix-a-parroting-chatbot-and",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T02:48:18.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/kharlampiiev/fine-tuning-qwen25-15b-on-a-mac-to-fix-a-parroting-chatbot-and-halve-the-prompt-23j9"
  },
  "original_language": "en",
  "account": "Three identical responses to the same question – \"could you tell me a joke?\" – presented a problem. Small language models often repeat earlier turns in the conversation instead of following the rules in a long system prompt. The first step to fix the issue was to adjust the prompt and code. A 91-token reminder was added to the last user turn, an off-topic classifier was implemented, and history cleanup was performed. However, the prompt was lengthy and slow on two cores. The plan then was to fine-tune the model's behavior while retaining facts in the prompt. A 1.5B model trained on facts was chosen, as it tends to blend facts and invent related ones when trained on other data. The fine-tuned model was trained on generated data using a locally hosted Qwen2.5-14B-Instruct model, with the facts-only system prompt taken directly from the production code. This resulted in 828 records in raw.jsonl format, which were then manually reviewed and filtered. Out of the 686 accepted records, 617 were used for training and 69 for validation in the mlx-lm chat format. The model was then fine-tuned using LoRA with MLX, learning only from the assistant's replies, with a top 16 of 28 layers. The trainable parameters accounted for just 0.684% of the total model size. After converting the fine-tuned model to GGUF format, it achieved a validation loss of 0.745 at its best and 0.78 at the end of training. The tuned model performed better than the previous version, declining jokes, refusing to answer about ATS or HubSpot, and responding in Russian when asked about building sites. The speed remained relatively unchanged, with only a slight increase in the prompt size.",
  "summary": "By Georgii Kharlampiiev, Mindscend Our company website is a single chat page backed by a model we host ourselves. This post walks through how we fine-tuned that model with LoRA on a MacBook, how we checked it was better, and how we shipped it to a CPU-only server. All the numbers are from our own runs. The stack Model: Qwen2.5-1.5B-Instruct, GGUF, Q4_K_M (about 1 GB). Serving: llama.cpp…",
  "key_points": [
    "Fine-tuned Qwen2.5-1.5B model to reduce parroting in chatbot",
    "Added 91-token reminder to last user turn and implemented off-topic classifier",
    "Achieved 0.745 validation loss with 0.684% trainable parameters"
  ],
  "editors_take": "Fine-tuning the Qwen2.5-1.5B model on a Mac resolved a parroting issue in a chatbot, allowing it to follow conversation rules while reducing the prompt length by half.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}