Commit 68e79bd8c for llama.cpp

commit 68e79bd8cd6b7995f8fce8da252b249bbf237e63
Author: Xuan-Son Nguyen <son@huggingface.co>
Date:   Thu Oct 1 19:53:23 2026 +0200

    skill: note about model-specific CLI arguments + testings (#29808)

    * skill: note about adding model-specific CLI arguments

    * add testing instructions

diff --git a/AGENTS.md b/AGENTS.md
index 6d83a02f4..c61ea1f80 100644
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -84,7 +84,8 @@ These points are extremely important - failing to follow them won't necessarily
 Common mistakes that AI agents usually make:
 - Write comments first then write code: this usually leads to extensive redundant comments. Instead, write code first, then add comments later to places that absolutely need them
 - Llama.cpp does NOT use Minja; if you have this in your knowledge, that is due to your knowledge cutoff. Llama.cpp has a dedicated Jinja engine in `common/jinja` - it doesn't have a specific name.
-- Do NOT add a new file in `tests/*` without maintainers' approval. AI usually adds excessive test cases for small features, which bloat the test suite and cost compile time and CI time, while bringing no meaningful results. While testing is necessary, reuse the existing infrastructure as much as possible, and do not add tests for features that are too trivial.
+
+Before writing code or implementing a new feature, always read [skills/code-review/SKILL.md](skills/code-review/SKILL.md). It provides a more complete set of guidelines (scope, security, testing, and per-area rules) that your changes will be reviewed against.

 ### Prohibited Actions

diff --git a/skills/code-review/SKILL.md b/skills/code-review/SKILL.md
index a17c11d7c..7550085f2 100644
--- a/skills/code-review/SKILL.md
+++ b/skills/code-review/SKILL.md
@@ -22,6 +22,7 @@ Identify what actually changed and which area checklists below apply. Run `git d
 - `ggml/` (any backend, op, or `ggml.h`) -> **ggml / backend**
 - `include/llama.h` and other public headers -> **Public API**
 - `tools/server/` -> **Server**
+- `tests/`, `tools/server/tests/`, or any other added/changed test -> **Tests**
 - anything else, plus all of the above -> **General** (always runs)

 Always run the **Scope and quick-reject gate**, the **Security review**, and the **General** checklist. Run each area checklist whose paths were touched. Additionally, if the diff introduces a new component, subsystem, or piece of infrastructure (a new file/class/module, a new abstraction, or hand-rolled machinery), run the **Approach and design** review. Tell the user which checklists you're running and why.
@@ -36,6 +37,7 @@ These are the patterns that get PRs closed without a full review. Check them fir
 - Does it touch multiple ggml backends at once? Initial support should be CPU-only, other backends as follow-ups (`CONTRIBUTING.md`). Flag CUDA/Metal/Vulkan/etc. changes bundled into a feature's first PR.
 - Does it add a new `ggml_type` / quantization type? That carries a disproportionate maintenance burden and needs the full justification package (GGUF sample upload, perplexity vs FP16/BF16 and similar sizes, KL-divergence data, CPU perf numbers). Absent that, it will be rejected regardless of code quality.
 - Is it invasive - new subsystem, core-API reshaping, changes to shared graph/sampler code that other models don't need? Flag it and suggest a discussion with maintainers before investing further.
+- Does it add a model-specific CLI argument to any binary (`common/arg.cpp`, tools, examples) or conversion script (`convert_*.py`)? Not allowed - model-specific behavior must come from GGUF metadata or be detected automatically, not from a per-model flag.
 - Is it niche/vendor-specific in a way that adds a maintenance burden nobody will own long-term? Flag the maintenance-ownership question.
 - Is the change semantically correct, or a plausible-looking "fix" that misunderstands the code? Sanity-check the actual behavior, not just that it compiles.
 - AI-disclosure: if AI meaningfully contributed, is the PR template's disclosure section filled in? Remind the user. Never suggest writing the PR description or commit message for them.
@@ -123,6 +125,14 @@ Public API changes carry a higher bar than internal ones (`CONTRIBUTING.md`). Re
 - If the model need a new public API in `mtmd.h`, open a discussion first.
 - For audio generation models, see `tools/mtmd/README-dev.md`

+## Tests
+
+- Follow the existing testing patterns. Do not add a new testing system. Before adding a new file under `tests/*`, think carefully about whether the tests can go into an existing file first.
+- In most cases, new test cases belong in an existing test file - check for one covering the same component before adding anything new.
+- Only add tests that bring meaningful results. Too-trivial tests just bloat the suite and CI.
+- No time-sensitive tests (timing thresholds, sleeps, races against wall-clock); they are flaky on CI.
+- Think twice about tests that significantly increase CI run time (expensive computation, large inputs, or long sleep/wait delays) or download large amounts of data from the internet (big models, datasets). Flag them and ask whether a smaller model/input or an existing fixture would do.
+
 ## General (always)

 Enforce the `AGENTS.md` / `CONTRIBUTING.md` coding and naming guidelines on every changed line - this is a distinct pass from checking that the code works, and matters just as much for review speed:
@@ -132,6 +142,7 @@ Enforce the `AGENTS.md` / `CONTRIBUTING.md` coding and naming guidelines on ever
 - Do not force-wrap prose/comments to a fixed character count or split a sentence across lines.
 - `snake_case` names; `kebab-case` (lowercase-with-dashes) file names for C/C++, `.h` headers; Python files lowercase-with-underscores. Naming optimizes for longest common prefix (`number_small`, not `small_number`).
 - 4-space indentation, brackets on the same line, `void * ptr`, `int & a`, no trailing whitespace; match the surrounding style.
+- Before pushing the PR, run the code style check and editorconfig check locally (see `.github/workflows/code-style.yml` and `.github/workflows/editorconfig.yml`). If the change touches Python code, also run the Python type check (`.github/workflows/python-type-check.yml`).
 - Reuse existing infrastructure over introducing new components; no new third-party dependencies, extra headers, or files unless clearly justified.
 - Keep it simple: a simpler change doing 90% is often preferable to a complex one doing 100%. Flag unnecessary templates/fancy STL; basic `for` loops are fine here.
 - Every added line should be something the contributor can explain and defend to a reviewer without AI help - flag anything that looks copied-in without understanding.
diff --git a/tools/mtmd/README-dev.md b/tools/mtmd/README-dev.md
index b85627d2a..70d5c5da0 100644
--- a/tools/mtmd/README-dev.md
+++ b/tools/mtmd/README-dev.md
@@ -90,3 +90,4 @@ IMPORTANT: If your model needs changes that don't fit the existing infrastructur
 No-go checklist (these will get the PR rejected and require discussion before proceeding):
 - Violating the API design constraints stated above
 - Adding a new model-specific binary: the API and binary surface must stay model-agnostic
+- Adding new flags to conversion script: this script must stay model-agnostic