Task recipes


Five recipes cover the task lifecycle: task_create_from_cloud.py creates one task from object keys already in a registered bucket, tasks_bulk_from_cloud.py creates a whole batch of tasks in a project from that same bucket, task_inspect_and_export.py inspects an existing task, exports its dataset locally, and reports analytics from its event log, tasks_create_per_label_group.py splits annotation work over one set of images into several tasks, one per shape type, and task_create_job_mapping.py creates a task with an explicit file-to-job mapping.

Create a task from cloud object keys

Creates a task from images that already live in a registered bucket.

Flag Required Meaning
--host yes Server URL
--token yes Personal Access Token
--cloud-storage-id yes Registered cloud storage id (see cloud_storage_register.py)
--cloud-keys yes Object keys in the bucket, space-separated
--name no Task name (default 'Task from cloud storage')
--labels no Label names, space-separated (default object)
--cleanup no Delete the created task at the end
python task_create_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \
    --cloud-storage-id 7 --cloud-keys 'images/0001.jpg' 'images/0002.jpg' \
    --labels car person

The script

# Copyright (C) CVAT.ai Corporation
#
# SPDX-License-Identifier: MIT

"""Create an annotation task from images that already live in a registered
cloud storage.

Steps:
  1. Create a task whose data is a list of object keys in the bucket.
  2. Print the result.
  3. Optionally delete it (--cleanup).

Register a bucket first with cloud_storage_register.py to get the storage id.

Usage (run ``python task_create_from_cloud.py --help`` for the full list of options):
  python task_create_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --cloud-storage-id 7 --cloud-keys 'images/0001.jpg' 'images/0002.jpg' \\
      --labels car person
"""

import argparse

from cvat_sdk import make_client, models
from cvat_sdk.core.proxies.tasks import ResourceType


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
    parser.add_argument("--host", required=True, help="CVAT server URL, e.g. 'https://app.cvat.ai'")
    parser.add_argument(
        "--token",
        required=True,
        help="Personal Access Token (CVAT UI: Profile -> Security)",
    )
    parser.add_argument(
        "--cloud-storage-id",
        type=int,
        required=True,
        help="a registered cloud storage id (see cloud_storage_register.py)",
    )
    parser.add_argument(
        "--cloud-keys",
        nargs="+",
        required=True,
        help="object keys in the bucket, e.g. 'images/0001.jpg' 'images/0002.jpg'",
    )
    parser.add_argument(
        "--name",
        default="Task from cloud storage",
        help="task name (default: '%(default)s')",
    )
    parser.add_argument(
        "--labels", nargs="+", default=["object"], help="label names (default: %(default)s)"
    )
    parser.add_argument("--cleanup", action="store_true", help="delete the created task at the end")
    return parser.parse_args()


def main() -> None:
    args = parse_args()
    with make_client(args.host, access_token=args.token) as client:
        # ResourceType.SHARE + cloud_storage_id = read images from the bucket
        task = client.tasks.create_from_data(
            spec=models.TaskWriteRequest(
                name=args.name,
                labels=[models.PatchedLabelRequest(name=name) for name in args.labels],
            ),
            resource_type=ResourceType.SHARE,
            resources=args.cloud_keys,
            data_params={"cloud_storage_id": args.cloud_storage_id},
        )
        print(f"Created task {task.id} with {task.size} frames: {args.host}/tasks/{task.id}")

        if args.cleanup:
            task.remove()
            print(f"Deleted task {task.id}")
        else:
            print("Keeping the task; pass --cleanup to delete it")


if __name__ == "__main__":
    main()

Bulk-create tasks in a project from a bucket

Creates several tasks in one call, all inside the same project, each reading its data from a registered cloud storage. Two ways to spell a task’s data, repeatable and mixable: --task KEY[,KEY,...] lists explicit object keys (a single key makes a video/single-image task; multiple keys make an image task whose frames are those keys in order), and --task-pattern PATTERN makes one task from every bucket file matching a fnmatch wildcard (e.g. 'batch_a/*.jpg'), resolved from the bucket’s manifest instead of listing every key by hand. Because every task belongs to the project, they share its label schema — no --labels here.

Flag Required Meaning
--host yes Server URL
--token yes Personal Access Token
--cloud-storage-id yes Registered cloud storage id (see cloud_storage_register.py)
--project-id yes Project the tasks are created in; supplies the labels
--task KEY[,KEY,...] one of --task / --task-pattern One --task per task; repeat the flag for more
--task-pattern PATTERN one of --task / --task-pattern One task per wildcard, matched via the bucket’s manifest; repeat for more
--manifest no Manifest object key used to resolve --task-pattern (default 'manifest.jsonl')
--name-prefix no Task-name prefix; each task is named <prefix> N (default 'Bulk task')
--cleanup no Delete every created task at the end
# three video tasks in project 42
python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \
    --cloud-storage-id 7 --project-id 42 \
    --task 'videos/clip_01.mp4' --task 'videos/clip_02.mp4' --task 'videos/clip_03.mp4'

# two image-batch tasks in project 42
python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \
    --cloud-storage-id 7 --project-id 42 \
    --task 'batch_a/img_1.jpg,batch_a/img_2.jpg' \
    --task 'batch_b/img_1.jpg,batch_b/img_2.jpg'

# the same two batches, without listing every key: one task per wildcard match
python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \
    --cloud-storage-id 7 --project-id 42 --manifest manifest.jsonl \
    --task-pattern 'batch_a/*.jpg' --task-pattern 'batch_b/*.jpg'

The script

# Copyright (C) CVAT.ai Corporation
#
# SPDX-License-Identifier: MIT

"""Bulk-create tasks inside a project, each task's data read from a registered
cloud storage.

Two ways to spell a task's data, repeatable and mixable:
  --task KEY[,KEY,...]    explicit object keys, in order:
                             * a single key -> a video task (or single-image task);
                             * several keys -> an image task, in the given order.
  --task-pattern PATTERN  every bucket file matching a fnmatch wildcard (e.g.
                           'batch_a/*.jpg'), resolved from the bucket's
                           manifest instead of being listed one by one.

All tasks land in the same project, so they share its label schema.

Steps:
  1. For each --task, create a task in --project-id from its explicit keys.
  2. For each --task-pattern, create a task in --project-id from every bucket
     file the wildcard matches, resolved via the bucket's manifest.
  3. Print the created ids and a summary count.
  4. Optionally delete every created task (--cleanup).

Register a bucket first with cloud_storage_register.py to get the storage id.
A --task-pattern also needs a manifest file already generated for the bucket -
see "How to generate manifest file" in the CVAT docs on attaching cloud storage.

Usage (run ``python tasks_bulk_from_cloud.py --help`` for the full list of options):
  # three video tasks in project 42
  python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --cloud-storage-id 7 --project-id 42 \\
      --task 'videos/clip_01.mp4' --task 'videos/clip_02.mp4' --task 'videos/clip_03.mp4'

  # two image-batch tasks in project 42
  python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --cloud-storage-id 7 --project-id 42 \\
      --task 'batch_a/img_1.jpg,batch_a/img_2.jpg' \\
      --task 'batch_b/img_1.jpg,batch_b/img_2.jpg'

  # the same two batches, without listing every key: one task per wildcard match
  python tasks_bulk_from_cloud.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --cloud-storage-id 7 --project-id 42 --manifest manifest.jsonl \\
      --task-pattern 'batch_a/*.jpg' --task-pattern 'batch_b/*.jpg'
"""

import argparse

from cvat_sdk import make_client, models
from cvat_sdk.core.proxies.tasks import ResourceType


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
    parser.add_argument("--host", required=True, help="CVAT server URL, e.g. 'https://app.cvat.ai'")
    parser.add_argument(
        "--token",
        required=True,
        help="Personal Access Token (CVAT UI: Profile -> Security)",
    )
    parser.add_argument(
        "--cloud-storage-id",
        type=int,
        required=True,
        help="a registered cloud storage id (see cloud_storage_register.py)",
    )
    parser.add_argument(
        "--project-id",
        type=int,
        required=True,
        help="tasks are created in this project and inherit its labels",
    )
    parser.add_argument(
        "--task",
        dest="tasks",
        action="append",
        default=[],
        metavar="KEY[,KEY,...]",
        help="comma-separated object keys for one task; repeat for more tasks",
    )
    parser.add_argument(
        "--task-pattern",
        dest="task_patterns",
        action="append",
        default=[],
        metavar="PATTERN",
        help="one task from every bucket file matching this fnmatch wildcard "
        "(e.g. 'batch_a/*.jpg'); repeat for more tasks. Needs --manifest. "
        "(default: '%(default)s')",
    )
    parser.add_argument(
        "--manifest",
        default="manifest.jsonl",
        help="manifest object key in the bucket, used to resolve --task-pattern "
        "(default: '%(default)s')",
    )
    parser.add_argument(
        "--name-prefix",
        default="Bulk task",
        help="task name prefix; each task is named '<prefix> N' (default: '%(default)s')",
    )
    parser.add_argument(
        "--cleanup", action="store_true", help="delete every created task at the end"
    )
    args = parser.parse_args()
    if not args.tasks and not args.task_patterns:
        parser.error("at least one --task or --task-pattern is required")
    return args


def main() -> None:
    args = parse_args()
    task_key_groups = [
        [key.strip() for key in spec.split(",") if key.strip()] for spec in args.tasks
    ]
    if any(not group for group in task_key_groups):
        raise SystemExit("each --task must contain at least one non-empty key")

    with make_client(args.host, access_token=args.token) as client:
        created = []
        for keys in task_key_groups:
            # Tasks in a project inherit the project's labels — do NOT pass labels.
            # ResourceType.SHARE + cloud_storage_id reads the objects from the bucket.
            task = client.tasks.create_from_data(
                spec=models.TaskWriteRequest(
                    name=f"{args.name_prefix} {len(created) + 1}", project_id=args.project_id
                ),
                resource_type=ResourceType.SHARE,
                resources=keys,
                data_params={"cloud_storage_id": args.cloud_storage_id},
            )
            created.append(task)
            print(f"Created task {task.id} ({task.size} frames): {args.host}/tasks/{task.id}")

        for pattern in args.task_patterns:
            # A wildcard task needs the bucket's manifest as its only resource;
            # the server expands filename_pattern against it (fnmatch syntax).
            # use_cache=True is required to serve data straight from the bucket.
            task = client.tasks.create_from_data(
                spec=models.TaskWriteRequest(
                    name=f"{args.name_prefix} {len(created) + 1}", project_id=args.project_id
                ),
                resource_type=ResourceType.SHARE,
                resources=[args.manifest],
                data_params={
                    "cloud_storage_id": args.cloud_storage_id,
                    "use_cache": True,
                    "filename_pattern": pattern,
                },
            )
            created.append(task)
            print(
                f"Created task {task.id} ({task.size} frames) from pattern {pattern!r}: "
                f"{args.host}/tasks/{task.id}"
            )

        print(f"Created {len(created)} tasks in project {args.project_id}")

        if args.cleanup:
            for task in created:
                task.remove()
            print(f"Deleted {len(created)} tasks")
        else:
            print("Keeping the tasks; pass --cleanup to delete them")


if __name__ == "__main__":
    main()

Inspect a task and export its dataset

Prints a summary of an existing task (labels, jobs, frames), exports its dataset to a local zip, then exports the task’s event log and reports two analytics computed from it: how many people are currently assigned to a job, and how many jobs were rejected in review and sent back for rework.

Flag Required Meaning
--host yes Server URL
--token yes Personal Access Token
--task-id yes Id of the task to inspect and export
--export-format no Exporter name (default 'COCO 1.0')
python task_inspect_and_export.py --host 'https://app.cvat.ai' --token '<your token>' \
    --task-id 42 --export-format 'COCO 1.0'

The script

# Copyright (C) CVAT.ai Corporation
#
# SPDX-License-Identifier: MIT

"""Inspect an existing task (labels, jobs, frames), export its dataset to a
local zip, and export its event log to report quick analytics.

Steps:
  1. Retrieve the task and print a summary: labels, jobs (stage/state), frames.
  2. Fetch the server's export format list and validate --export-format.
  3. Export the dataset to task_<id>_dataset.zip in the current directory.
  4. Export the task's event log to task_<id>_events.csv and report two
     analytics: how many people are currently assigned to a job, and how
     many jobs were rejected in review and sent back for rework - the second
     one needs the log, since a job's current state doesn't show its history.

Usage (run ``python task_inspect_and_export.py --help`` for the full list of options):
  python task_inspect_and_export.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --task-id 42 --export-format 'COCO 1.0'
"""

import argparse
import csv
import sys
from pathlib import Path

from cvat_sdk import make_client
from cvat_sdk.core.downloading import Downloader
from cvat_sdk.core.proxies.types import Location


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
    parser.add_argument("--host", required=True, help="CVAT server URL, e.g. 'https://app.cvat.ai'")
    parser.add_argument(
        "--token",
        required=True,
        help="Personal Access Token (CVAT UI: Profile -> Security)",
    )
    parser.add_argument(
        "--task-id", type=int, required=True, help="id of an existing task, e.g. 42"
    )
    parser.add_argument(
        "--export-format",
        default="COCO 1.0",
        help="exporter name, e.g. 'COCO 1.0' (default: '%(default)s')",
    )
    return parser.parse_args()


def count_reworks(events_path: Path) -> int:
    """Count how many times a job in the log was rejected in review, i.e. sent
    back to the annotator for rework. A job's current state only shows where
    it stands now, not how many times it got there, so this needs the log.
    """
    with events_path.open(newline="") as f:
        return sum(
            1
            for row in csv.DictReader(f)
            if row["scope"] == "update:job"
            and row["obj_name"] == "state"
            and row["obj_val"] == "rejected"
        )


def main() -> None:
    args = parse_args()
    with make_client(args.host, access_token=args.token) as client:
        # 1. Inspect
        task = client.tasks.retrieve(args.task_id)
        jobs = task.get_jobs()
        print(f"Task {task.id}: {task.name!r}, {task.size} frames")
        print(f"  labels: {[label.name for label in task.get_labels()]}")
        for job in jobs:
            print(f"  job {job.id}: stage={job.stage}, state={job.state}")

        # 2. Validate the export format against the server's list.
        # Low-level API: there is no high-level proxy for the format list yet.
        formats, _ = client.api_client.server_api.retrieve_annotation_formats()
        names = [f.name for f in formats.exporters]
        if args.export_format not in names:
            sys.exit(
                f"Unknown export format {args.export_format!r}. Choose one of: {', '.join(names)}"
            )

        # 3. Export the dataset to a local zip
        local_path = Path(f"task_{task.id}_dataset.zip")
        task.export_dataset(
            args.export_format, local_path, include_images=False, location=Location.LOCAL
        )
        print(f"Exported {local_path.resolve()}")

        # 4. Export the task's event log and report quick analytics.
        events_path = Path(f"task_{task.id}_events.csv")
        Downloader(client).prepare_and_download_file_from_endpoint(
            client.api_client.events_api.create_export_endpoint,
            events_path,
            query_params={"task_id": task.id},
        )
        print(f"Exported {events_path.resolve()}")

        assigned = {job.assignee.id for job in jobs if job.assignee}
        print(f"  {len(assigned)} people currently assigned, {count_reworks(events_path)} reworks")


if __name__ == "__main__":
    main()

Split work into one task per label group

Creates one task per --task 'NAME:TYPE:label1,label2' spec over the same images, with that spec’s labels typed to its shape type. Boxes, polygons, and tags then get annotated in parallel, each in a task that shows only the labels its annotator needs.

The tasks are standalone, not tasks of one project: tasks in a project share the project’s label set, so a per-task label set cannot exist inside a single project — the SDK rejects labels on a task that has a project_id. Every spec is validated before anything is created, and --cleanup deletes the tasks created so far even if a later one fails.

Flag Required Meaning
--host yes Server URL
--token yes Personal Access Token
--image-dir yes Directory with the images to annotate; every file in it is uploaded
--task NAME:TYPE:LABELS yes One task; repeat for more
--segment-size no Frames per job, in every task
--cleanup no Delete the created tasks at the end

TYPE is one of rectangle, polygon, polyline, points, ellipse, cuboid, mask, tag, any.

python tasks_create_per_label_group.py --host 'https://app.cvat.ai' --token '<your token>' \
    --image-dir ./images \
    --task 'boxes:rectangle:car,person' \
    --task 'roads:polygon:road,lane' \
    --task 'weather:tag:rain,snow'

The script

# Copyright (C) CVAT.ai Corporation
#
# SPDX-License-Identifier: MIT

"""Create one task per shape type and label group over the same set of images,
so boxes, polygons, and tags are annotated in parallel by different people
instead of all at once in one crowded task.

Each --task spec is 'NAME:TYPE:label1,label2': the task's name, the shape type
its labels are drawn with, and the labels themselves.

The tasks are standalone, not tasks of one project: tasks in a project share the
project's label set, so a per-task label set cannot exist inside a single
project.

Steps:
  1. Parse and validate every --task spec before creating anything.
  2. Collect the files from --image-dir.
  3. Create one task per spec, with that spec's labels typed to its shape type.
  4. Print a summary of what each task got.

Usage (run ``python tasks_create_per_label_group.py --help`` for the full list of options):
  python tasks_create_per_label_group.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --image-dir ./images \\
      --task 'boxes:rectangle:car,person' \\
      --task 'roads:polygon:road,lane' \\
      --task 'weather:tag:rain,snow'
"""

import argparse
import sys
from pathlib import Path

from cvat_sdk import make_client, models
from cvat_sdk.core.proxies.tasks import ResourceType

LABEL_TYPES = {
    "any",
    "cuboid",
    "ellipse",
    "mask",
    "points",
    "polygon",
    "polyline",
    "rectangle",
    "tag",
}


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
    parser.add_argument("--host", required=True, help="CVAT server URL, e.g. 'https://app.cvat.ai'")
    parser.add_argument(
        "--token",
        required=True,
        help="Personal Access Token (CVAT UI: Profile -> Security)",
    )
    parser.add_argument(
        "--image-dir",
        type=Path,
        required=True,
        help="directory with the images to annotate; every file in it is uploaded, "
        "and the server decides which media it accepts",
    )
    parser.add_argument(
        "--task",
        action="append",
        required=True,
        metavar="NAME:TYPE:LABELS",
        help="one task, e.g. 'boxes:rectangle:car,person' (repeat for more)",
    )
    parser.add_argument("--segment-size", type=int, help="frames per job, in every task")
    parser.add_argument(
        "--cleanup", action="store_true", help="delete the created tasks at the end"
    )
    return parser.parse_args()


def parse_spec(spec: str) -> tuple[str, str, list[str]]:
    """'boxes:rectangle:car,person' -> ('boxes', 'rectangle', ['car', 'person'])."""
    name, _, rest = spec.partition(":")
    type_, _, labels = rest.partition(":")
    names = [label.strip() for label in labels.split(",") if label.strip()]
    if not (name and type_ and names):
        sys.exit(f"Bad --task {spec!r}: expected 'NAME:TYPE:label1,label2'")
    if type_ not in LABEL_TYPES:
        sys.exit(
            f"Unknown label type {type_!r} in --task {spec!r}. "
            f"Choose one of: {', '.join(sorted(LABEL_TYPES))}"
        )
    return name, type_, names


def main() -> None:
    args = parse_args()
    # 1. Validate every spec first, so a typo in the last one costs nothing.
    specs = [parse_spec(spec) for spec in args.task]

    # 2. The same files go into every task. They are passed as they are found:
    # the server is the authority on which media formats it supports, so
    # filtering by extension here would only reject files CVAT can read.
    images = sorted(p for p in args.image_dir.iterdir() if p.is_file())
    if not images:
        sys.exit(f"No files found in {args.image_dir}")

    created = []
    with make_client(args.host, access_token=args.token) as client:
        try:
            for name, type_, label_names in specs:
                task = client.tasks.create_from_data(
                    spec=models.TaskWriteRequest(
                        name=name,
                        labels=[
                            models.PatchedLabelRequest(name=label, type=type_)
                            for label in label_names
                        ],
                        **({"segment_size": args.segment_size} if args.segment_size else {}),
                    ),
                    resource_type=ResourceType.LOCAL,
                    resources=images,
                )
                created.append(task)
                print(
                    f"Created task {task.id} {name!r} "
                    f"({type_}: {', '.join(label_names)}, {len(task.get_jobs())} job(s)): "
                    f"{args.host}/tasks/{task.id}"
                )
            print(f"Created {len(created)} task(s) from {len(images)} file(s)")
        finally:
            # Clean up whatever was created, including after a mid-run failure.
            if args.cleanup:
                for task in created:
                    task.remove()
                    print(f"Deleted task {task.id}")
            elif created:
                print("Keeping the tasks; pass --cleanup to delete them")


if __name__ == "__main__":
    main()

Decide which files go into which job

Normally CVAT cuts a task into jobs of segment_size frames. The job_file_mapping data parameter replaces that with an explicit grouping: one job per camera, per scene, or per delivery batch. Group the files yourself with repeated --job flags, or chunk the directory with --files-per-job N — a count of files, the explicit equivalent of segment_size.

Every file in --image-dir must belong to exactly one job — unknown files, duplicates, and leftovers are rejected before the task is created. Afterwards the recipe reads the jobs back from the server and writes job_file_mapping.csv, one job_id,frame,file_name row per file: the mapping you see is the one that exists, and every file names the frame it actually landed on. Frame numbers are zero-based indexes within the task.

The example creates one object label so the task is ready for annotation.

job_file_mapping implies predefined file ordering and takes one file per frame, so it applies to image tasks only: a video is a single file the server cuts into frames itself, and there is nothing to map. CVAT rejects the request if the task’s data turns out to be a video.

Flag Required Meaning
--host yes Server URL
--token yes Personal Access Token
--image-dir yes Directory with the task’s images, one file per frame
--job FILE [FILE ...] one of --job / --files-per-job The files of one job; repeat for more
--files-per-job N one of --job / --files-per-job Number of files per job: chunk the directory into jobs of N files each
--name no Task name
--output no Mapping CSV path (default job_file_mapping.csv)
--cleanup no Delete the created task at the end
python task_create_job_mapping.py --host 'https://app.cvat.ai' --token '<your token>' \
    --image-dir ./images \
    --job 'cam1_001.png' 'cam1_002.png' --job 'cam2_001.png' 'cam2_002.png'

The script

# Copyright (C) CVAT.ai Corporation
#
# SPDX-License-Identifier: MIT

"""Create a task whose jobs are defined by you, file by file, instead of by a
frame count: one job per camera, per scene, per delivery batch.

Two ways to group:
  --job FILE [FILE ...]   one job per occurrence, in the given file order
  --files-per-job N       a count, not a file list: chunk the sorted directory
                          listing into jobs of N files each, the explicit
                          equivalent of the task's segment_size

Steps:
  1. Build the file groups and check them against --image-dir: every file must
     exist, appear once, and belong to a job.
  2. Create the task with the job_file_mapping data parameter.
  3. Read the jobs back from the server and print/write the resulting mapping,
     so what you see is what the server built, not what was requested.

job_file_mapping implies predefined file ordering and takes one file per frame,
so it applies to image tasks only: a video is a single file the server cuts into
frames itself, and there is nothing to map. A directory of images is therefore
what this recipe takes, and CVAT rejects the request if the data turns out to be
a video.

Usage (run ``python task_create_job_mapping.py --help`` for the full list of options):
  python task_create_job_mapping.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --image-dir ./images \\
      --job 'cam1_001.png' 'cam1_002.png' --job 'cam2_001.png' 'cam2_002.png'
  python task_create_job_mapping.py --host 'https://app.cvat.ai' --token '<your token>' \\
      --image-dir ./images --files-per-job 50
"""

import argparse
import csv
import sys
from pathlib import Path

from cvat_sdk import make_client, models
from cvat_sdk.core.proxies.tasks import ResourceType


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
    parser.add_argument("--host", required=True, help="CVAT server URL, e.g. 'https://app.cvat.ai'")
    parser.add_argument(
        "--token",
        required=True,
        help="Personal Access Token (CVAT UI: Profile -> Security)",
    )
    parser.add_argument(
        "--image-dir",
        type=Path,
        required=True,
        help="directory with the task's images, one file per frame (a video cannot be "
        "mapped to jobs file by file)",
    )
    grouping = parser.add_mutually_exclusive_group(required=True)
    grouping.add_argument(
        "--job",
        action="append",
        nargs="+",
        metavar="FILE",
        help="the files of one job (repeat for more jobs)",
    )
    grouping.add_argument(
        "--files-per-job",
        type=int,
        metavar="N",
        help="number of files per job: chunk the sorted directory listing into jobs of N "
        "files each (the explicit equivalent of the task's segment_size)",
    )
    parser.add_argument("--name", default="Task with a job file mapping", help="task name")
    parser.add_argument(
        "--output",
        type=Path,
        default=Path("job_file_mapping.csv"),
        help="path to write the resulting mapping to (default: %(default)s)",
    )
    parser.add_argument("--cleanup", action="store_true", help="delete the created task at the end")
    return parser.parse_args()


def build_groups(args: argparse.Namespace, available: list[str]) -> list[list[str]]:
    """The file groups to send as job_file_mapping, validated against the directory."""
    if args.files_per_job:
        if args.files_per_job < 1:
            sys.exit("--files-per-job must be at least 1")
        return [
            available[start : start + args.files_per_job]
            for start in range(0, len(available), args.files_per_job)
        ]

    groups = [list(group) for group in args.job]
    known = set(available)
    assigned = []
    for group in groups:
        assigned.extend(group)

    unknown = [name for name in assigned if name not in known]
    if unknown:
        sys.exit(f"File(s) {', '.join(unknown)} not found in {args.image_dir}")

    duplicates = sorted({name for name in assigned if assigned.count(name) > 1})
    if duplicates:
        sys.exit(f"File(s) {', '.join(duplicates)} appear in more than one job")

    unassigned = sorted(known - set(assigned))
    if unassigned:
        sys.exit(
            f"File(s) {', '.join(unassigned)} are not assigned to any job. "
            "Every file in --image-dir must belong to exactly one --job."
        )
    return groups


def main() -> None:
    args = parse_args()
    # The files are taken as they are found: the server is the authority on
    # which media formats it supports, so filtering by extension here would
    # only reject files CVAT can read.
    available = sorted(p.name for p in args.image_dir.iterdir() if p.is_file())
    if not available:
        sys.exit(f"No files found in {args.image_dir}")

    groups = build_groups(args, available)
    resources = [args.image_dir / name for group in groups for name in group]
    print(f"Requesting {len(groups)} job(s) over {len(resources)} file(s)")

    with make_client(args.host, access_token=args.token) as client:
        task = client.tasks.create_from_data(
            spec=models.TaskWriteRequest(
                name=args.name,
                labels=[models.PatchedLabelRequest(name="object")],
            ),
            resource_type=ResourceType.LOCAL,
            resources=resources,
            data_params={"job_file_mapping": groups},
        )
        print(f"Created task {task.id} with {task.size} frames: {args.host}/tasks/{task.id}")

        # 3. The mapping is built using one row per file, carrying
        # the frame that file ended up on. A job's frames come back in order,
        # so the frame is the job's start plus the file's position in it.
        rows = []
        for job in sorted(task.get_jobs(), key=lambda job: job.start_frame):
            names = [frame.name for frame in job.get_frames_info()]
            print(f"  job {job.id} frames {job.start_frame}-{job.stop_frame}: {', '.join(names)}")
            rows.extend(
                {"job_id": job.id, "frame": job.start_frame + offset, "file_name": name}
                for offset, name in enumerate(names)
            )

        with args.output.open("w", newline="") as f:
            writer = csv.DictWriter(f, fieldnames=["job_id", "frame", "file_name"])
            writer.writeheader()
            writer.writerows(rows)
        print(f"Wrote {args.output.resolve()}")

        if args.cleanup:
            task.remove()
            print(f"Deleted task {task.id}")
        else:
            print("Keeping the task; pass --cleanup to delete it")


if __name__ == "__main__":
    main()

Other SDK options:

SDK method / parameter What it adds
client.tasks.create_from_data(..., resource_type=ResourceType.LOCAL | SHARE | REMOTE) Where resources come from: LOCAL (upload local files), SHARE (keys in a cloud storage / mounted share), REMOTE (URLs). Defaults to LOCAL.
client.tasks.create_from_data(..., data_params={...}) Extra data options as a dict, e.g. image_quality (1-100), sorting_method ("lexicographical"/"natural"/"predefined"/"random"), cloud_storage_id (int).
client.tasks.create_from_data(..., annotation_path="path.zip", annotation_format="CVAT XML 1.1") Upload an initial annotations file at creation. annotation_path is a str file path; annotation_format is a str, default "CVAT XML 1.1".
client.tasks.create_from_data(..., status_check_period=<int seconds>, pbar=ProgressReporter()) status_check_period (int, seconds) is the upload status poll interval (defaults to Config.status_check_period); pbar is a cvat_sdk.core.progress.ProgressReporter for upload progress.
client.tasks.list(..., search=, sort=) Free-text search and server-side ordering (sort), in addition to filter.
client.tasks.create_from_backup(path) Recreate a task from a task backup archive.
Task.import_annotations(format_name, path) Load annotations into an existing task - the import counterpart of export_dataset.
Task.get_frame(frame_id: int, *, quality="original" | "compressed") Return a single frame as a file-like object (io.RawIOBase) of image bytes. quality is an optional keyword argument ("original" or "compressed"); if omitted, the server default is used.
Task.download_frames(frame_ids: Sequence[int], outdir=".", quality="original", image_extension=None, filename_pattern="frame_{frame_id:06d}{frame_ext}") Save the given frames to disk under outdir. image_extension (e.g. "png") overrides the auto-detected extension; quality is "original" or "compressed".
Task.get_meta() / Task.get_frames_info() Read frame count, chunk layout, and per-frame metadata.
Task.export_dataset(..., pbar=ProgressReporter()) Report local-download progress (a cvat_sdk.core.progress.ProgressReporter).
Task.export_dataset(..., status_check_period=<int seconds>) Poll interval (int, seconds) between server status checks; defaults to Config.status_check_period.
Task.export_dataset(filename=<directory>) Pass a directory as filename for a local export and the server-generated file name is used.
Task.export_dataset(..., location=Location.CLOUD_STORAGE, cloud_storage_id=<int>) Export straight to a registered cloud storage instead of downloading locally.
client.api_client.events_api.create_export(project_id=, job_id=, user_id=, _from=, to=) Scope or time-bound the event-log export beyond a single task.
data_params={"job_file_mapping": [[...], [...]]} Define each job’s files explicitly instead of using segment_size.
PatchedLabelRequest(name=..., type="polygon") Restrict a label to one shape type, as the per-label-group recipe does.
Job.get_frames_info() The frames a job really holds, names included.

Notes: