A Day With a 1.5B Model: What Qwen2.5-Coder Can (and Can't) Do
A hands-on account of running Qwen2.5-Coder-1.5B as the workhorse of a real data pipeline, with the numbers to back it up.
There is a lot of marketing enthusiasm around small local models, and less careful accounting of what happens when you put one to work on a real pipeline, unattended, for hours at a time. This is that accounting, for Qwen2.5-Coder-1.5B-Instruct — a Q4_K_M GGUF quantization running locally via llama.cpp — used as the first-pass generator in a pipeline whose end goal is a Drupal-specific dataset accurate enough to help engineers migrate Drupal 7 sites to modern versions faster.
The role it was given
The model's job: read a source document (an API reference page, or a question-and-answer pair from a technical forum), and produce a structured instruction/response JSON object grounded in that source — without copying it verbatim, and without inventing facts not present in it. It's a demanding task for a 1.5B model, and every improvement along the way had to be earned against real, measured data rather than a hand-picked example.
Discovery #1: it has strong instincts — sometimes too strong
The most interesting behavior we found: when the input context was long or noisy, the model would occasionally reach for a confident, plausible answer from its training data rather than the specific source text in front of it. We caught the same well-formed, memorized answer appearing across documentation pages about entirely different classes — a clear signal that the model had a strong prior about "the usual answer" and needed a nudge toward the specific text at hand.
The fix was direct: instead of a generic "don't make things up" instruction, we named the exact recurring content explicitly and told the model plainly not to reach for it unless the source genuinely discussed it. This worked immediately and durably — a precise, specific instruction outperformed a broad one by a wide margin.
Discovery #2: it takes your worked example very literally
We included a worked example (few-shot) in the prompt to demonstrate the expected output format — standard, well-supported practice. At volume, the example's fictional names occasionally reappeared verbatim in real outputs, a pattern serious enough to require a dedicated fix (the exact measured leak rate, and how we closed it, is covered in a companion post on few-shot prompting). The short version: once we understood why it was happening — a plausible-sounding example is hard for a small model to distinguish from real content — the fix was straightforward: swap the example for something deliberately unmistakable, and name the leak explicitly as off-limits. This lines up with published research on "demonstration regurgitation" in small models, which was a useful sanity check that we'd diagnosed the actual mechanism rather than just patched a symptom.
Discovery #3: give it a checklist, not just an instruction
When we required the output to contain both an exact copied signature and a plain-language explanation of behavior, the model sometimes returned only the signature. Rather than fight this in prose, we moved the explanation into its own required field in a JSON Schema, enforced through grammar-constrained decoding at the inference layer. This resolved the issue completely, because in this case the model already knew what a good explanation looked like — it just needed a structural nudge to always include one.
Discovery #4: a reasonable idea to test — and the discipline to revert it
We tried one more idea: instead of letting the model search for and copy a method signature from source code, we handed it a pre-verified, numbered list of valid options and asked it to choose one — removing even the possibility of copying a nonexistent signature. It's a reasonable hypothesis, and worth testing. In this case, measured live rather than on a small hand-checked sample, the first-pass acceptance rate for that pipeline moved from roughly 40% to 21.4%, because the model would pick a valid item and then describe a different, memorized one instead.
We reverted, kept the numbered list as a reference for the model rather than a rigid mechanism, and moved on within the same working session. That willingness to test an idea, measure it honestly against live traffic rather than a favorable sample, and revert when the data says no is what made the rest of the day's progress trustworthy.
What actually moved the needle
In a controlled, same-verifier, same-sample comparison (8 real documents):
Baseline (temperature 0.2)
1/8 accepted
Self-consistency (3 samples, majority vote)
0/8
Best-of-3 with heuristic scoring
0/8
Temperature 0 (fully greedy decoding)
2/8 — best result
No few-shot example
0/8
Rotating between 3 few-shot examples
0/8
Multiple-choice symbol selection
0/8
Greedy decoding produced the best result in this comparison — moving from 1/8 to 2/8. We want to be precise about what that is: a sample of 8 is small enough that "doubled" is a preliminary signal, not a settled result. It was consistent with a separate, larger-scale observation over a subsequent unattended run (discussed below), which is what gave us enough confidence to keep it as the production default — but the 8-item comparison alone shouldn't be read as statistically conclusive on its own.
The bottom line, in numbers
Across a 7-hour unattended production run, the model achieved a 46.0% first-pass acceptance rate on the documentation-generation task (176 verified pairs written) and 26.4% on the harder, less-structured community-Q&A task (53 verified pairs) — up from a 0% starting point on that second task earlier in the project, after a separate, unrelated fix to an overly strict upstream filter. The 46.0% figure matches the best rate this pipeline has ever recorded, historically, with a noticeably healthier mix of remaining edge cases.
None of this makes a 1.5B model a substitute for a larger one. What it does show is that a small model, engineered around its specific and fairly predictable failure modes, can carry the bulk of a high-volume generation workload at effectively zero marginal cost — which matters directly for the underlying goal: producing a large enough, accurate enough Drupal dataset to be genuinely useful for migration work, without the cost of running every single generation through an expensive model.