How-to guides

Run Nomad jobs from inline Python or airflow-config YAML. Both formats use the Nomad CLI on the Airflow worker to reach the cluster’s API.

How to select a cluster and namespace

Install the Nomad CLI on every worker that can execute this lifecycle. Set the connection environment in the worker process, rather than only in an interactive shell:

export NOMAD_ADDR=https://nomad.example.com:4646
export NOMAD_NAMESPACE=analytics
export NOMAD_CACERT=/etc/nomad/tls/ca.pem

Supply NOMAD_TOKEN through your deployment’s secret injection. If the cluster requires mutual TLS, also supply NOMAD_CLIENT_CERT and NOMAD_CLIENT_KEY. The token needs permissions to read, submit, and stop jobs in the target namespace. See the Nomad CLI reference for connection and authentication variables.

Create the namespace before submitting jobs. Have a cluster administrator run:

nomad namespace apply analytics

Set job.namespace="analytics" in Python or namespace: analytics under job in YAML. This explicit namespace is included in the jobspec and takes precedence for job-scoped CLI calls. If omitted, the CLI uses NOMAD_NAMESPACE or default.

Confirm access from the worker account:

nomad job status -namespace=analytics

Select datacenters and drivers available in that cluster. Commands and paths in a task’s config refer to the Nomad client node or task container, rather than the Airflow worker. The exec examples below need Linux Nomad clients with that driver available. No SSH connection is required for a remote cluster.

How to run a job in Python

Install airflow-nomad[airflow] in Airflow 2, or airflow-nomad[airflow3] in Airflow 3. Complete the cluster setup.

Give all lifecycle tasks the same writable absolute working_dir. With several workers, mount that path on shared storage; configuration, registration, and cleanup must reach the same jobspec. Every worker also needs the same cluster credentials. Reserve the job ID for this DAG.

Save this DAG in the DAG folder, adapting the namespace, datacenter, driver, and command to your cluster:

from datetime import datetime, timezone

from airflow import DAG
from airflow_nomad import Job, Nomad, NomadAirflowConfiguration, Task, TaskGroup

with DAG(
    dag_id="batch-nomad",
    schedule="@daily",
    start_date=datetime(2025, 1, 1, tzinfo=timezone.utc),
    catchup=False,
) as dag:
    nomad = Nomad(
        dag=dag,
        cfg=NomadAirflowConfiguration(
            working_dir="/var/tmp/batch-nomad",
            job=Job(
                id="airflow-batch",
                type="batch",
                namespace="analytics",
                datacenters=["dc1"],
                task_groups=[
                    TaskGroup(
                        name="batch",
                        tasks=[Task(name="sleep", driver="exec", config={"command": "/bin/sleep", "args": ["5"]})],
                    ),
                ],
            ),
        ),
    )

Run airflow tasks list batch-nomad, then trigger batch-nomad. Registration starts the job; successful completion stops it and removes the local jobspec.

How to run a job with airflow-config

Use the packages, cluster, and shared artifact path from the Python guide. Install airflow-config in the DAG parsing environment.

Create config/batch_nomad.yaml beside the DAG loader:

# @package _global_
_target_: airflow_config.Configuration
_convert_: all

dags:
  batch-nomad:
    schedule: "@daily"
    start_date: "2025-01-01"
    catchup: false
    tasks:
      run:
        _target_: airflow_nomad.NomadTask
        cfg:
          working_dir: /var/tmp/batch-nomad
          job:
            id: airflow-batch
            type: batch
            namespace: analytics
            datacenters: [dc1]
            task_groups:
              - name: batch
                tasks:
                  - name: sleep
                    driver: exec
                    config:
                      command: /bin/sleep
                      args: ["5"]

Save batch_nomad.py in the DAG folder:

"""Generate Airflow DAGs from the Nomad configuration."""

from airflow_config import load_config

config = load_config("config", "batch_nomad")
config.generate_in_mem()

_convert_: all converts nested command arguments into Python lists before model construction, including values inside Nomad’s untyped config mapping.

Run airflow tasks list batch-nomad, then trigger batch-nomad. Deploy either this loader or the inline Python DAG for that DAG ID. Environment variables select local or remote clusters in either format.

How to limit monitoring or retain a service

Set monitoring controls under cfg in YAML:

check_interval: 00:00:10
check_timeout: 08:00:00
runtime: 04:00:00
maxretrigger: 3

In Python, pass timedelta(seconds=10), timedelta(hours=8), and timedelta(hours=4) for the duration fields, importing timedelta from datetime. runtime or endtime ends monitoring through the stop branch; check_timeout is the sensor timeout. See the timing reference for reference dates.

For a service that should remain running between DAG runs, set stop_on_exit=False, cleanup=False, restart_on_initial=True, and restart_on_retrigger=True in Python, or under cfg in YAML:

stop_on_exit: false
cleanup: false
restart_on_initial: true
restart_on_retrigger: true
runtime: 04:00:00

Set job.type="service" and configure the long-running command or container in its task group. Keep the job ID, namespace, and artifact paths stable. Set an end condition for monitoring a service that never exits. Configure Nomad restart and reschedule policies separately for failures between Airflow runs.

How to purge job history or clean up a failed run

To purge history when the normal stop task runs, set purge_on_exit=True in Python, or under cfg in YAML:

stop_on_exit: true
purge_on_exit: true
cleanup: true

Stop and cleanup follow successful monitoring. They do not finalize every failure. After a failed or cancelled run, stop the job from an authorized Nomad CLI environment, using its actual namespace and ID:

nomad job stop -namespace=analytics -purge -yes airflow-batch

Remove its generated JSON file only after no active run uses it. The generated force-kill task also requests a purge; its normal parent is skipped, so it is not an automatic failure cleanup path.

How to chain the lifecycle with other tasks

In Python, connect existing Airflow tasks to the Nomad object:

prepare >> nomad >> publish

In airflow-config, use the task mapping key in dependencies:

tasks:
  prepare:
    _target_: airflow_pydantic.BashTask
    bash_command: echo prepare
  run:
    _target_: airflow_nomad.NomadTask
    _convert_: all
    dependencies: [prepare]
    cfg:
      working_dir: /var/tmp/chained-nomad
      job:
        id: airflow-chained
        type: batch
        namespace: analytics
        datacenters: [dc1]
        task_groups:
          - name: batch
            tasks:
              - name: sleep
                driver: exec
                config:
                  command: /bin/sleep
                  args: ["5"]
  publish:
    _target_: airflow_pydantic.BashTask
    bash_command: echo publish
    dependencies: [run]

Upstream tasks precede configuration; downstream tasks follow cleanup. Keep stop and cleanup enabled for downstream tasks with the default all_success trigger rule. Disabling cleanup creates a skipped boundary task.