Support `auto_doctring` in Processors #42101

yonigozlan · 2025-11-07T17:31:21Z

What does this PR do?

Add support for Processors in @auto_docstring, and many other improvements to auto_docstring.py and check_docstrings.py, including more robust auto-fix with check_docstrings for missing, redundant, or unnecessary docstrings.

For processors, auto_docstring will pull custom args docstrings from custom "Kwargs" TypeDicts and add them to the .__doc__. For example, for processing_aria, we have:

class AriaImagesKwargs(ImagesKwargs, total=False):
    """
    split_image (`bool`, *optional*, defaults to `False`):
        Whether to split large images into multiple crops. When enabled, images exceeding the maximum size are
        divided into overlapping crops that are processed separately and then combined. This allows processing
        of very high-resolution images that exceed the model's input size limits.
    max_image_size (`int`, *optional*, defaults to `980`):
        Maximum image size (in pixels) for a single image crop. Images larger than this will be split into
        multiple crops when `split_image=True`, or resized if splitting is disabled. This parameter controls
        the maximum resolution of individual image patches processed by the model.
    min_image_size (`int`, *optional*):
        Minimum image size (in pixels) for a single image crop. Images smaller than this will be upscaled to
        meet the minimum requirement. If not specified, images are processed at their original size (subject
        to the maximum size constraint).
    """

    split_image: bool
    max_image_size: int
    min_image_size: int


class AriaProcessorKwargs(ProcessingKwargs, total=False):
    images_kwargs: AriaImagesKwargs

    _defaults = {
        "text_kwargs": {
            "padding": False,
            "return_mm_token_type_ids": False,
        },
        "images_kwargs": {
            "max_image_size": 980,
            "split_image": False,
        },
        "return_tensors": TensorType.PYTORCH,
    }


@auto_docstring
class AriaProcessor(ProcessorMixin):
    ...

    @auto_docstring
    def __call__(
        self,
        text: Union[TextInput, PreTokenizedInput, list[TextInput], list[PreTokenizedInput]],
        images: Optional[ImageInput] = None,
        **kwargs: Unpack[AriaProcessorKwargs],
    ) -> BatchFeature:
        r"""
        Returns:
            [`BatchFeature`]: A [`BatchFeature`] with the following fields:
            - **input_ids** -- List of token ids to be fed to a model. Returned when `text` is not `None`.
            - **attention_mask** -- List of indices specifying which tokens should be attended to by the model (when
            `return_attention_mask=True` or if *"attention_mask"* is in `self.model_input_names` and if `text` is not
            `None`).
            - **pixel_values** -- Pixel values to be fed to a model. Returned when `images` is not `None`.
            - **pixel_mask** -- Pixel mask to be fed to a model. Returned when `images` is not `None`.
        """
        ...

which results in the following docstring:

print(AriaProcessor.__call__.__doc__)

        Args:
            text (`Union[str, list, list]`):
                The sequence or batch of sequences to be encoded. Each sequence can be a string or a list of strings
                (pretokenized string). If the sequences are provided as list of strings (pretokenized), you must set
                `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).
            images (`Union[PIL.Image.Image, numpy.ndarray, torch.Tensor, list, list, list]`, *optional*):
                Image to preprocess. Expects a single or batch of images with pixel values ranging from 0 to 255. If
                passing in images with pixel values between 0 and 1, set `do_rescale=False`.
            split_image (`bool`, *optional*, defaults to `False`):
                Whether to split large images into multiple crops. When enabled, images exceeding the maximum size are
                divided into overlapping crops that are processed separately and then combined. This allows processing
                of very high-resolution images that exceed the model's input size limits.
            max_image_size (`int`, *optional*, defaults to `980`):
                Maximum image size (in pixels) for a single image crop. Images larger than this will be split into
                multiple crops when `split_image=True`, or resized if splitting is disabled. This parameter controls
                the maximum resolution of individual image patches processed by the model.
            min_image_size (`int`, *optional*):
                Minimum image size (in pixels) for a single image crop. Images smaller than this will be upscaled to
                meet the minimum requirement. If not specified, images are processed at their original size (subject
                to the maximum size constraint).
            return_tensors (`str` or [`~utils.TensorType`], *optional*):
                If set, will return tensors of a particular framework. Acceptable values are:

                - `'pt'`: Return PyTorch `torch.Tensor` objects.
                - `'np'`: Return NumPy `np.ndarray` objects.
        Returns:
            [`BatchFeature`]: A [`BatchFeature`] with the following fields:
            - **input_ids** -- List of token ids to be fed to a model. Returned when `text` is not `None`.
            - **attention_mask** -- List of indices specifying which tokens should be attended to by the model (when
            `return_attention_mask=True` or if *"attention_mask"* is in `self.model_input_names` and if `text` is not
            `None`).
            - **pixel_values** -- Pixel values to be fed to a model. Returned when `images` is not `None`.
            - **pixel_mask** -- Pixel mask to be fed to a model. Returned when `images` is not `None`.

…asses

…rom-processors

… (temporarily)

…rom-processors

…m/yonigozlan/transformers into remove-attributes-from-processors

…rom-processors

* Super * Super * Super * Super --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

* detectron2 - part 1 * detectron2 - part 2 --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

…ingface#41978) fix autoawq[kernels] Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

…moving redundant docstring and placeholders

…ng-in-processor

HuggingFaceDocBuilderDev · 2026-01-06T16:45:17Z

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

github-actions · 2026-01-06T18:53:29Z

[For maintainers] Suggested jobs to run (before merge)

run-slow: align, altclip, aria, aya_vision, bamba, bark, blip, blip_2, bridgetower, bros, chameleon, chinese_clip, clap, clip, clipseg, clvp

yonigozlan · 2026-01-06T20:23:43Z

Also Cc @stevhliu :)

Cyrilvallez

Trusting you on that, but I think it would be time to add some proper tests no? I see a very old test_auto_docstrings.py but that does not run any tests -> probably a very nice idea to start rewriting it!

Cyrilvallez · 2026-01-07T09:40:04Z

src/transformers/utils/auto_docstring.py

+    intro = f"""Constructs a {class_name} which wraps {components_text} into a single processor.
+
+[`{class_name}`] offers all the functionalities of {classes_text}. See the
+{classes_text_short} for more information.
+"""


nit: can we use textwrap.dedent here, so that the string respects the function indentation?

Yep it's done right after

Humm, I don't see it 😅 I meant doing something like

intro = textwrap.dedent( """ bla bla more bla """ ).strip()

so that the indentation stays inside the function

yonigozlan · 2026-01-07T17:40:09Z

Trusting you on that, but I think it would be time to add some proper tests no? I see a very old test_auto_docstrings.py but that does not run any tests -> probably a very nice idea to start rewriting it!

Yes clearly! I'll add tests in the next autodocstring PR ;)

…ng-in-processor

stevhliu

nice, thanks! added a few nits to the parameter definitions :)

stevhliu · 2026-01-07T18:19:03Z

src/transformers/utils/auto_docstring.py

+
+    chat_template = {
+        "description": """
+    A Jinja template which will be used to convert lists of messages in a chat into a tokenizable string.


Suggested change

A Jinja template which will be used to convert lists of messages in a chat into a tokenizable string.

A Jinja template to convert lists of messages in a chat into a tokenizable string.

stevhliu · 2026-01-07T18:22:00Z

src/transformers/utils/auto_docstring.py

+    The sequence or batch of sequences to be encoded. Each sequence can be a string or a list of strings
+    (pretokenized string). If the sequences are provided as list of strings (pretokenized), you must set
+    `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).


Suggested change

The sequence or batch of sequences to be encoded. Each sequence can be a string or a list of strings

(pretokenized string). If the sequences are provided as list of strings (pretokenized), you must set

`is_split_into_words=True` (to lift the ambiguity with a batch of sequences).

The sequence or batch of sequences to be encoded. Each sequence can be a string or a list of strings

(pretokenized string). If you pass a pretokenized input, set `is_split_into_words=True` to avoid ambiguity with batched inputs.

stevhliu · 2026-01-07T18:24:40Z

src/transformers/utils/auto_docstring.py

+    """,
+    }
+
+    audio = {


just curious, whats the difference between audio and audios below it?

I think audios is deprecated but still present in some places

stevhliu · 2026-01-07T18:26:58Z

src/transformers/utils/auto_docstring.py

+    pad_to_multiple_of = {
+        "description": """
+    If set will pad the sequence to a multiple of the provided value. Requires `padding` to be activated.
+    This is especially useful to enable the use of Tensor Cores on NVIDIA hardware with compute capability


Suggested change

This is especially useful to enable the use of Tensor Cores on NVIDIA hardware with compute capability

This is especially useful to enable using Tensor Cores on NVIDIA hardware with compute capability

stevhliu · 2026-01-07T18:29:24Z

src/transformers/utils/auto_docstring.py

+    add_special_tokens = {
+        "description": """
+    Whether or not to add special tokens when encoding the sequences. This will use the underlying
+    `PretrainedTokenizerBase.build_inputs_with_special_tokens` function, which defines which tokens are


Suggested change

`PretrainedTokenizerBase.build_inputs_with_special_tokens` function, which defines which tokens are

[`PretrainedTokenizerBase.build_inputs_with_special_tokens`] function, which defines which tokens are

stevhliu · 2026-01-07T18:30:23Z

src/transformers/utils/auto_docstring.py

+    list of strings (pretokenized string). If the sequences are provided as list of strings (pretokenized),
+    you must set `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).


Suggested change

list of strings (pretokenized string). If the sequences are provided as list of strings (pretokenized),

you must set `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).

list of strings (pretokenized string). If you pass pretokenized input, set is_split_into_words=True to avoid ambiguity with batched inputs.

stevhliu · 2026-01-07T18:30:44Z

src/transformers/utils/auto_docstring.py

+    list of strings (pretokenized string). If the sequences are provided as list of strings (pretokenized),
+    you must set `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).


Suggested change

list of strings (pretokenized string). If the sequences are provided as list of strings (pretokenized),

you must set `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).

list of strings (pretokenized string). If you pass pretokenized input, set is_split_into_words=True to avoid ambiguity with batched inputs.

…ng-in-processor

…om/yonigozlan/transformers into support-auto_doctring-in-processor

…ng-in-processor

yonigozlan and others added 30 commits October 15, 2025 15:47

remove attributes and add all missing sub processors to their auto cl…

f48a47b

…asses

remove all mentions of .attributes

d5d5c58

cleanup

dd505b5

fix processor tests

6a1448f

fix modular

a292900

remove last attributes

63a255d

fixup

ef73759

Merge remote-tracking branch 'upstream/main' into remove-attributes-f…

b5e8b2e

…rom-processors

fixes after merge

f14ff3c

fix wrong tokenizer in auto florence2

0306430

fix missing audio_processor + nits

01cb815

Override __init__ in NewProcessor and change hf-internal-testing-repo…

49ec906

… (temporarily)

Merge remote-tracking branch 'upstream/main' into remove-attributes-f…

7dd5682

…rom-processors

fix auto tokenizer test

946cc5c

add init to markup_lm

b0cb3e0

update CustomProcessor in custom_processing

3b9e846

remove print

53de7a4

Merge branch 'main' into remove-attributes-from-processors

93d2c4d

Merge remote-tracking branch 'upstream/main' into remove-attributes-f…

feeec28

…rom-processors

nit

4a6b080

Merge branch 'remove-attributes-from-processors' of https://github.co…

02402a0

…m/yonigozlan/transformers into remove-attributes-from-processors

fix test modeling owlv2

757e1f1

fix test_processing_layoutxlm

bf763b2

Fix owlv2, wav2vec2, markuplm, voxtral issues

0799a0a

Merge remote-tracking branch 'upstream/main' into remove-attributes-f…

bf1a4b6

…rom-processors

add support for loading and saving multiple tokenizer natively

e3f130d

remove exclude_attributes from save_pretrained

cc45a7e

Run slow v2 (huggingface#41914)

6b9e7c9

* Super * Super * Super * Super --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

Fix detectron2 installation in docker files (huggingface#41975)

0ccb0e3

* detectron2 - part 1 * detectron2 - part 2 --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

Fix autoawq[kernels] installation in quantization docker file (hugg…

1eeece5

…ingface#41978) fix autoawq[kernels] Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>

yonigozlan added 3 commits January 6, 2026 16:21

Add recurring processor args to auto_docstring and add support for re…

68f178b

…moving redundant docstring and placeholders

replace placeholders with real docstrings

1ce14e9

Merge remote-tracking branch 'upstream/main' into support-auto_doctri…

0ee2c3f

…ng-in-processor

yonigozlan added 3 commits January 6, 2026 17:31

fix copies

22b29b8

fixup

8d5ffa8

remove unwanted changes

ab1f03b

yonigozlan changed the title ~~[WIP] Support auto_doctring in Processors~~ Support auto_doctring in Processors Jan 6, 2026

yonigozlan added 3 commits January 6, 2026 18:19

fix unprotected imports

525804c

Fix unprotected imports

852b458

fix unprotected imports

03d1cd3

Add __call__ to all docs of processors

22721cd

yonigozlan requested review from ArthurZucker and Cyrilvallez January 6, 2026 20:10

Cyrilvallez approved these changes Jan 7, 2026

View reviewed changes

Merge remote-tracking branch 'upstream/main' into support-auto_doctri…

b170599

…ng-in-processor

stevhliu approved these changes Jan 7, 2026

View reviewed changes

yonigozlan and others added 2 commits January 7, 2026 20:41

nits docs

b73220d

Merge branch 'main' into support-auto_doctring-in-processor

dcea25a

yonigozlan enabled auto-merge (squash) January 7, 2026 20:41

yonigozlan and others added 5 commits January 7, 2026 19:49

Merge branch 'main' into support-auto_doctring-in-processor

80b849f

Merge remote-tracking branch 'upstream/main' into support-auto_doctri…

edae136

…ng-in-processor

Merge branch 'support-auto_doctring-in-processor' of https://github.c…

14a5070

…om/yonigozlan/transformers into support-auto_doctring-in-processor

add flaky test

b3bf0e3

Merge remote-tracking branch 'upstream/main' into support-auto_doctri…

d639cd9

…ng-in-processor

yonigozlan merged commit c8bc4de into huggingface:main Jan 8, 2026
25 checks passed

vasqu mentioned this pull request Jan 8, 2026

[Timm] Increase tol in flaky test #43173

Closed

	A Jinja template which will be used to convert lists of messages in a chat into a tokenizable string.
	A Jinja template to convert lists of messages in a chat into a tokenizable string.

	This is especially useful to enable the use of Tensor Cores on NVIDIA hardware with compute capability
	This is especially useful to enable using Tensor Cores on NVIDIA hardware with compute capability

	`PretrainedTokenizerBase.build_inputs_with_special_tokens` function, which defines which tokens are
	[`PretrainedTokenizerBase.build_inputs_with_special_tokens`] function, which defines which tokens are

		list of strings (pretokenized string). If the sequences are provided as list of strings (pretokenized),
		you must set `is_split_into_words=True` (to lift the ambiguity with a batch of sequences).

Support auto_doctring in Processors #42101

Support auto_doctring in Processors #42101

Uh oh!

Conversation

yonigozlan commented Nov 7, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What does this PR do?

Uh oh!

HuggingFaceDocBuilderDev commented Jan 6, 2026

Uh oh!

github-actions bot commented Jan 6, 2026

Uh oh!

yonigozlan commented Jan 6, 2026

Uh oh!

Cyrilvallez left a comment

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Cyrilvallez Jan 8, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Choose a reason for hiding this comment

Uh oh!

yonigozlan commented Jan 7, 2026

Uh oh!

stevhliu left a comment

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

20 participants

Support `auto_doctring` in Processors #42101

Support `auto_doctring` in Processors #42101

yonigozlan commented Nov 7, 2025 •

edited

Loading

Cyrilvallez Jan 8, 2026 •

edited

Loading