Batch

AWS Batch — compute environments, job queues, job definitions, scheduling policies, consumable resources, service environments, quota shares, and the job control plane. restJson1 protocol.

AWS Batch (the batch service) runs batch computing workloads: you register a job definition, point a job queue at one or more compute environments, and submit jobs that run as containers.

The wedge: against every other free local emulator AWS Batch is a fake — its compute does no real work. MiniStack's Batch jumps a job straight to SUCCEEDED with no container; Moto runs Docker but leaks; LocalStack gates Batch behind its Ultimate tier. fakecloud already runs ECS tasks as real containers, and Batch is built to run real jobs on that same engine.

Supported today

  • Compute environments — CreateComputeEnvironment, DescribeComputeEnvironments, UpdateComputeEnvironment, DeleteComputeEnvironment. Created VALID / ENABLED.
  • Job queues — CreateJobQueue, DescribeJobQueues, UpdateJobQueue, DeleteJobQueue, with computeEnvironmentOrder, priority, and an optional schedulingPolicyArn.
  • Job definitions — RegisterJobDefinition (monotonic per-name revision), DescribeJobDefinitions (filter by name / ARN / status), DeregisterJobDefinition (marks the revision INACTIVE).
  • Scheduling policies — CreateSchedulingPolicy, DescribeSchedulingPolicies, ListSchedulingPolicies, UpdateSchedulingPolicy, DeleteSchedulingPolicy (fair-share).
  • Jobs — real container execution — SubmitJob launches the job definition's containerProperties (image / command / vcpus / memory / environment, with this submit's containerOverrides applied) as a real container on fakecloud's ECS task engine, and drives the job status off the container's actual lifecycle: SUBMITTED → STARTING → RUNNING → SUCCEEDED when the container exits 0, or FAILED (carrying the real container.exitCode) on a non-zero exit. This is the wedge: every other free emulator fakes Batch compute (MiniStack jumps straight to SUCCEEDED with no container). With no container runtime available the job stays SUBMITTED honestly — never an auto-success. DescribeJobs / ListJobs (filter by queue / status) report live status + exit code; CancelJob / TerminateJob stop a job.
  • Array jobs — SubmitJob with arrayProperties.size = N spawns N real child containers (<jobId>:<index>), each with AWS_BATCH_JOB_ARRAY_INDEX set so it can select its slice of work. The parent's status and arrayProperties.statusSummary aggregate the children live — SUCCEEDED only when every child exits 0.
  • Job dependencies — SubmitJob with dependsOn parks the job at PENDING and launches it only once every dependency has SUCCEEDED; if any dependency FAILED, the dependent job fails with "Dependent job failed". The wait never blocks the SubmitJob call.
  • Retry + timeout — retryStrategy.attempts re-launches a failed container up to that many times (each prior attempt recorded under attempts[]); timeout.attemptDurationSeconds caps each attempt and fails the job with "Job attempt duration exceeded timeout" if the container overruns.
  • Consumable resources: CreateConsumableResource, DescribeConsumableResource, ListConsumableResources (CONSUMABLE_RESOURCE_NAME filter with trailing * prefix match), UpdateConsumableResource (SET / ADD / REMOVE, clientToken replay applied once), DeleteConsumableResource, and ListJobsByConsumableResource (JOB_STATUS / JOB_NAME filters). Jobs declare requirements via the job definition's consumableResourceProperties or SubmitJob's consumableResourcePropertiesOverride (unknown resources are rejected). Requirements gate dispatch for real: a job whose quantity doesn't fit waits at RUNNABLE until capacity returns. inUseQuantity / availableQuantity are computed from jobs holding the resource; a NON_REPLENISHABLE resource stays consumed once a job has started.
  • Service environments and service jobs: CreateServiceEnvironment / DescribeServiceEnvironments / UpdateServiceEnvironment / DeleteServiceEnvironment (SAGEMAKER_TRAINING, must be DISABLED and detached from every queue before delete). Job queues accept serviceEnvironmentOrder (and take the environments' jobQueueType). SubmitServiceJob validates the queue type and state, the JSON serviceRequestPayload, fair-share shareIdentifier and quota-share quotaShareName rules, and clientToken reuse; DescribeServiceJob, ListServiceJobs (status plus JOB_NAME / BEFORE_CREATED_AT / AFTER_CREATED_AT / SHARE_IDENTIFIER / QUOTA_SHARE_NAME filters), UpdateServiceJob (schedulingPriority) and TerminateServiceJob round-trip. fakecloud's SageMaker has no training executor, so a service job is accepted into the queue and waits at RUNNABLE until terminated. It never reports a fabricated training run.
  • Quota shares: CreateQuotaShare, DescribeQuotaShare, ListQuotaShares, UpdateQuotaShare, DeleteQuotaShare (must be DISABLED; deleting terminates the share's remaining service jobs).
  • Queue snapshot: GetJobQueueSnapshot reports the RUNNABLE front of the queue in dispatch order, the first job per quota share, and capacity utilization of dispatched jobs (overall, per fair-share identifier, and per quota share).
  • Tags — TagResource, UntagResource, ListTagsForResource, on every Batch ARN including consumable resources, service environments, quota shares and service jobs.
  • CloudFormation — AWS::Batch::ComputeEnvironment, AWS::Batch::JobQueue, and AWS::Batch::JobDefinition are provisioned into the Batch control plane when a stack is created (and removed on stack delete). The provisioned resources persist across a restart in persistent mode.

Terraform / CloudFormation can provision a full Batch stack (aws_batch_compute_environment, aws_batch_job_queue, aws_batch_job_definition, aws_batch_scheduling_policy) and an SDK client can submit jobs (single, array, dependency-chained, retried, or timed-out) that run real containers and report their real exit codes. The terraform-provider-aws Batch acceptance suite (TestAccBatchComputeEnvironment_basic, TestAccBatchJobQueue_basic, TestAccBatchJobDefinition_basic) runs against fakecloud.