Building a Drupal Training Dataset: A Practical Playbook
A field report on building a Drupal-specific instruction dataset from scratch — pipeline design, quality gates, and the real numbers behind a working system.
Most write-ups about building instruction datasets for LLM fine-tuning or retrieval-augmented generation assume you already have clean, structured source material. Real projects rarely do — and that gap is where the interesting engineering happens. This post walks through the pipeline we built to turn scattered Drupal technical content into a structured, verified instruction dataset, and shares the numbers from a real production run.
Why: the actual problem this is solving
Drupal 7 has been out of official support for a while, but it still runs a large share of live production sites. Moving those sites to the modern Drupal 9/10/11 line is slow, expensive, technical work — the API surface changed substantially, and a lot of what a migration engineer needs is scattered across reference documentation that isn't written as a guide and years of community Q&A of uneven, sometimes version-mismatched quality. The goal of this dataset isn't dataset engineering for its own sake: it's building the knowledge base that lets an AI system meaningfully assist with reading legacy Drupal 7 code, explaining what it does, and proposing the modern equivalent — the specific, expensive bottleneck in most Drupal migrations.
Why Drupal specifically is a hard, worthwhile target
Two properties of Drupal make this exercise both demanding and useful:
- Version fragmentation. Drupal 7 and Drupal 9/10/11 are close enough in name to be confused and far enough apart in API surface that content written for one needs careful, version-aware handling before it's safe to apply to the other. Getting this right is a precondition for the dataset being useful for migration work at all.
- Reference documentation isn't Q&A. Official API documentation describes classes, methods, and hooks — it doesn't come pre-packaged as "a developer asks X, here's the answer." Community question-and-answer content is closer to how developers actually phrase problems. Turning both into consistent instruction pairs is the core transformation this pipeline performs.
The pipeline
At a high level, the pipeline has three phases per source type:
- Extraction — pull raw pages/threads, tag them by Drupal version, and strip boilerplate (navigation chrome, cross-version link lists, reference counters from API doc pages). This cleanup step alone meaningfully improved downstream quality once we measured how much it mattered.
- Generation — a small, fast, local model turns each cleaned source into a candidate
(instruction, context, response)triple. - Verification — a second, stronger model checks the candidate against the source and the generation contract, catching anything that invents content, mismatches the declared Drupal version, or violates the output schema. If it fails, the next model up regenerates from scratch. Only what survives verification gets written to the dataset.
This generate-then-verify structure (a cascade, covered in its own post) is what let a small, inexpensive model carry most of the workload while still producing a dataset we could trust for a downstream task where correctness — not just fluency — is what actually matters.
Turning rejection logs into a roadmap
A naive pipeline treats "the model returned valid JSON" as success. We went further: every rejection was logged and categorized, turning the verification stage into a diagnostic tool rather than just a gate. Over the life of the project this produced 1,251 categorized first-pass rejections, and having them broken down by cause is what made systematic improvement possible:
Rejection reason — Share
- Vague or incomplete answer: 29.9%
- Fabricated content not in the source: 23.3%
- Malformed JSON / wrong output shape: 19.3%
- Off-topic or wrong symbol referenced: 13.6%
- Wrong Drupal version discussed: 7.0%
- Missing a required explanation field: 3.9%
- Empty response: 2.7%
Each category got its own targeted fix rather than a generic "try harder" prompt tweak, and each fix measurably shrank its category.
Version fidelity: the fix with the clearest volume impact
One of the more counter-intuitive lessons came from questioning an existing safeguard rather than adding a new one. An early version of the pipeline used a strict pre-filter that discarded any Drupal question that didn't explicitly state its version, on the theory that ambiguous version tagging would poison the dataset. Once we looked closely, most developers simply don't state "Drupal 7" when it's obvious from context — the filter was discarding good data by the bucketful while barely touching the real problem: content that actively contradicts its declared version.
Replacing "discard anything ambiguous" with "discard only genuine contradictions, checked against the actual technical content" produced two separate, worth-distinguishing improvements in the community Q&A pipeline:
- Total accepted output roughly tripled run-over-run, because far fewer good candidates were being thrown away before they even reached generation.
- The generation model's own first-pass acceptance rate rose from 0% to 26.4% in the equivalent measurement window, because the candidates that now reached it were, on average, less ambiguous and easier to get right.
These are two different metrics telling two different parts of the same story — one is about how much good source material makes it into the pipeline at all, the other is about how well the model performs once it gets there — and it's worth keeping them separate rather than treating either number as a stand-in for the other.
Results from a real run
After the improvements described above, a 7-hour unattended run produced:
- 176 verified pairs from the reference-documentation pipeline, at a 46.0% first-pass acceptance rate for the small generation model — matching the best rate ever recorded for this pipeline.
- 53 verified pairs from the community Q&A pipeline, at the 26.4% first-pass rate noted above.
Both pipelines are still improving, and a well-instrumented pipeline turns "the model got it wrong" from a dead end into the next concrete thing to fix — with the end goal, always, being a dataset accurate and complete enough to meaningfully speed up real Drupal 7 to Drupal 9/10/11 migration work, not a benchmark number for its own sake.
Takeaways
- Categorize rejections before you try to fix them — a precise breakdown turns "it fails sometimes" into a prioritized, tractable list.
- Question your existing safeguards, not just your new ideas. Our biggest single throughput win came from loosening an overcautious filter, not tightening one.
- Keep "how much data reaches your model" and "how well your model performs on what reaches it" as separate, separately-reported numbers. Conflating them makes a real improvement sound like a different, larger one than it is.