Danemar Parceros

We Open-Sourced the Dataset: drupal7-dataset-api on Hugging Face

Open-source Drupal 7 API dataset with 7,952 question-answer examples for fine-tuning, evaluation, and RAG. GPL-2.0-or-later, ready to download from Hugging Face.

If you've been following this series, you already know the shape of this story. In "A Day With a 1.5B Model: What Qwen2.5-Coder Can (and Can't) Do" we spent a day pushing a small local model against real Drupal 7 code and wrote down, honestly, where it held up and where it didn't. In "Building a Drupal Training Dataset: A Practical Playbook" we walked through the pipeline we built to turn that experimentation into something reusable: a small model drafting candidate question-answer pairs, stronger models acting as judges, and a lot of deliberate filtering before anything made it into the final set.

Today that dataset stops being an internal artifact. We've published it publicly on Hugging Face: drupal7-dataset-api. This post is the technical detail behind that release — what's in it, how each part was built, what we deliberately left out, and what we've already confirmed it's good for.

What it is

drupal7-dataset-api is 7,952 examples about the Drupal 7 API, extracted exclusively from api.drupal.org. It's meant for fine-tuning, for evaluating a model's Drupal 7 knowledge, or as a retrieval corpus for RAG. It ships in Parquet format, 8.77 MB, English only. It was created by Daniel Ricardo Ramirez Marin as part of Codicem, the internal project at Danemar Parceros behind this dataset pipeline. License is GPL-2.0-or-later, inherited directly from api.drupal.org — we didn't choose a different license than the source material carries.

The dataset ships as two configs, and they were built in almost opposite ways.

The reference config: 4,382 examples, LLM-generated and LLM-checked

reference holds questions about API symbols — a function, a constant, a class, a global variable — paired with its exact signature and a short explanation. The breakdown by symbol type:

  • 4,010 functions
  • 247 constants
  • 91 classes
  • 34 globals

This config was generated, not extracted. The pipeline worked in stages: several LLMs proposed candidate question/answer pairs from the source material on api.drupal.org; other LLMs acted as judges, discarding pairs that were fabricated or didn't hold up against the actual symbol; and a separate translation pass moved the content from Spanish to English, with explicit validation that code identifiers — variable names, function names, constants, quoted strings — survived the translation untouched. That last check matters more than it sounds: it's exactly the kind of step that's easy to skip and ends up quietly corrupting a dataset meant to teach a model precise API signatures.

The code config: 3,570 examples, zero LLM intervention

code is a different animal. It holds coding tasks where the answer is the real body of a Drupal 7 function — not a description of what the code does, but the code itself. Breakdown:

  • 2,501 functions
  • 739 hook implementations
  • 330 hook definition examples (from *.api.php files)

No LLM touched this config at any stage. It's extracted directly from the official Drupal 7 core source (7.104-dev). The tasks are generated from the docblocks that already exist above each function in the source; the responses are the exact function bodies, copied as-is. If reference is "can a model explain this API accurately," code is "can a model reproduce or recognize the actual implementation" — grounded in source code with no generation step to introduce drift.

What we excluded, and why

A dataset is as much about what you leave out as what you keep. During curation, these were excluded deliberately, with counts kept rather than quietly dropped:

  • 1,294 simpletest scaffolding examples — test boilerplate that doesn't teach API knowledge.
  • 35 explanations that were empty or too short to be useful.
  • 5 signatures that didn't match the source exactly — better to drop than ship a wrong signature.
  • 20 references to other Drupal versions that slipped in but don't belong in a Drupal 7-specific dataset.
  • 5 examples flagged as possible personal data.
  • 28 translations that failed the es→en identifier-preservation check described above.
  • 413 class methods with no reliable way to verify them against api.drupal.org.

That last number is worth sitting with. 413 examples were cut not because they were wrong, but because there was no way to confirm they were right. That's the kind of call that's easy to skip under deadline pressure and hard to notice later — if accuracy matters more than volume, you make it anyway.

Does it actually work?

We didn't just publish it and move on. We used drupal7-dataset-api to fine-tune a Qwen 7B model and confirmed a measurable improvement in its Drupal 7 knowledge afterward. The exact percentage? The dataset card doesn't report one, and neither do we — we're not going to round a number that isn't there. What we can say is that it's not a dataset sitting untested: it's already been run through a real fine-tuning pass with a real result.

Why this matters if you maintain or migrate Drupal 7

Drupal 7 reached official end of life on January 5, 2025 — no more official security patches for core or contrib. As of a BuiltWith snapshot taken September 2, 2026, there were still 149,752 active Drupal 7 sites identified globally. That's a lot of code that still needs people (and, increasingly, AI tooling) who can read it accurately, whether the goal is patching it, auditing it, or migrating it off Drupal 7 entirely. API knowledge is the first thing any of that work depends on — you can't safely rewrite or explain code you don't actually understand the surface of. That's the specific gap this dataset is built to help close.

How to use it

Each record carries: id, drupal_version (always "7"), source_type, source_url, instruction, context, response, tags, and created_at. That's enough structure to filter by source type, trace any example back to its exact page on api.drupal.org, or split the two configs differently than we did if your use case calls for it. It's on Hugging Face under the standard datasets loading path, ready for fine-tuning, benchmarking a model's Drupal 7 knowledge, or indexing as a RAG corpus.

License

GPL-2.0-or-later, inherited from api.drupal.org. If you build something on top of this dataset, that's the license your derivative work needs to respect.

This dataset came out of Codicem, the pipeline work we've written about in this series while building better tooling for the Drupal 7 → 11 migrations we do at Danemar Parceros. We're releasing it because better API knowledge for any model working on Drupal 7 code is useful regardless of who's using it — not just us.

drupal7-dataset-api on Hugging Face →