How to Structure Python IaC Projects for Scale

Scaling infrastructure requires deterministic execution paths and strict boundary enforcement. Ad-hoc scripts fail under multi-account, multi-region deployments. Treat infrastructure code as production software: typed, tested, and reviewed. This guide establishes architectural patterns that guarantee state integrity, secure credential handling, and fully testable Python IaC workflows, extending the conventions in IaC Design Principles.

1. Enforce Strict Directory Layouts & Module Boundaries

Isolation prevents configuration bleed and simplifies dependency resolution. Separate environment variables from infrastructure logic. Define explicit import paths for multi-cloud deployments. This architecture aligns with established patterns documented in Python IaC Fundamentals & Strategy.

Enforce Strict Directory Layouts & Module Boundaries Enforce Strict Directory Layouts & Module Boundaries: layered from pip down to SDK. pip Enforce Strict Directory Module Boundaries Python IaC Fundamentals SDK
Enforce Strict Directory Layouts & Module Boundaries: the building blocks this section assembles.

CLI: mkdir -p infra/{core,envs,tests,components} CLI: poetry init --no-interaction

Initialize the workspace with a locked dependency manager. Never rely on system-wide pip installations. Pin provider SDK versions explicitly.

Recommended layout:

infra/
  components/   # Reusable resource constructs (VPC, EKS, RDS)
  envs/         # Environment-specific overrides (dev.py, prod.py)
  core/         # State backend config, provider initialization
  tests/        # Unit and integration tests
__main__.py     # Pulumi entry point (imports from infra/)
cdktf.json      # CDKTF project config
pyproject.toml  # Dependencies and tooling config

The layout is only half of the boundary; the other half is a rule about which direction imports may flow. envs/ may import components/ and core/. components/ may import core/. Nothing may import envs/. Violating that last rule is how a "shared" component acquires a production-specific default, and it happens gradually enough that no single review catches it.

Make the rule executable rather than aspirational. import-linter reads contracts from pyproject.toml and fails the build when a layer reaches upward:

# CLI: lint-imports --config pyproject.toml
[tool.importlinter]
root_package = "infra"

[[tool.importlinter.contracts]]
name = "Environments are leaves"
type = "layers"
layers = ["infra.envs", "infra.components", "infra.core"]

The failure output names both ends of the illegal edge — infra.components.vpc is not allowed to import infra.envs.prod — which is far more actionable than discovering the coupling six months later when the dev environment cannot be built without production credentials.

One structural decision people defer too long: whether components/ stays in the repository or becomes an installed package. Keep it in-repo while the interfaces are still moving, because a version bump per change is friction you do not need. Extract it to a package once two or more repositories consume the same component, at which point the semantics of a breaking change become explicit rather than implicit. The mechanics of that extraction, including the JSII considerations for CDKTF, are covered in publishing CDKTF constructs as a Python package.

  • Verify __init__.py exports only public component APIs
  • Run python -m py_compile infra/ across all modules to catch syntax errors early
  • Assert no circular imports via pytest --import-mode=importlib
  • Add a lint-imports step to CI so the layer contract is enforced on every pull request

2. Implement Python 3.9+ Type Safety & Dependency Isolation

Dynamic typing introduces silent failures during stack synthesis. Enforce strict contracts at initialization. Validate configuration payloads before provider SDKs consume them. Immutable configuration patterns follow the guidelines in IaC Design Principles.

Implement Python + Type Safety & Dependency Isolation Implement Python + Type Safety & Dependency Isolation: Any then Dict then Implement Python then Type Safety then Dependency Any Dict Implement Python Type Safety Dependency
Implement Python + Type Safety & Dependency Isolation: the stages run left to right — Any, Dict, Implement Python, Type Safety, Dependency.

CLI: poetry add pulumi pulumi-aws cdktf pydantic CLI: mypy --strict --python-version 3.9 infra/

Lock your dependency tree to prevent provider SDK drift. Reject untyped Any or Dict in provider input signatures. Enforce from __future__ import annotations at the top of every module.

# infra/components/vpc.py
from __future__ import annotations

import pulumi
import pulumi_aws as aws
from dataclasses import dataclass
from typing import Sequence
from pydantic import BaseModel, IPvAnyNetwork, Field

class VpcConfig(BaseModel):
    cidr: IPvAnyNetwork
    public_subnets: Sequence[str]
    private_subnets: Sequence[str]
    enable_nat_gateway: bool = Field(default=True)

@dataclass(frozen=True)
class VpcOutputs:
    vpc_id: pulumi.Output[str]
    public_subnet_ids: pulumi.Output[Sequence[str]]
    private_subnet_ids: pulumi.Output[Sequence[str]]

def provision_vpc(config: VpcConfig, project_name: str) -> VpcOutputs:
    """Provision a strictly typed VPC with validated CIDR boundaries."""
    current_region = aws.get_region().name
    vpc = aws.ec2.Vpc(
        resource_name=f"{project_name}-main-vpc",
        cidr_block=str(config.cidr),
        enable_dns_support=True,
        enable_dns_hostnames=True,
    )

    public_subnets = [
        aws.ec2.Subnet(
            resource_name=f"{project_name}-pub-{idx}",
            vpc_id=vpc.id,
            cidr_block=cidr,
            map_public_ip_on_launch=True,
            availability_zone=f"{current_region}{'a' if idx == 0 else 'b'}",
        )
        for idx, cidr in enumerate(config.public_subnets)
    ]

    private_subnets = [
        aws.ec2.Subnet(
            resource_name=f"{project_name}-priv-{idx}",
            vpc_id=vpc.id,
            cidr_block=cidr,
            availability_zone=f"{current_region}{'a' if idx == 0 else 'b'}",
        )
        for idx, cidr in enumerate(config.private_subnets)
    ]

    return VpcOutputs(
        vpc_id=vpc.id,
        public_subnet_ids=pulumi.Output.all(*[s.id for s in public_subnets]),
        private_subnet_ids=pulumi.Output.all(*[s.id for s in private_subnets]),
    )
  • Run mypy infra/components/vpc.py with zero errors
  • Verify pydantic model rejects invalid CIDR blocks at runtime
  • Execute pytest tests/test_vpc.py to assert resource graph generation

3. Configure State Backends & Drift Detection Pipelines

Remote state requires distributed locking and encryption at rest. Never store state locally in CI or shared environments. Schedule automated drift scans to detect manual console modifications, following the convergence and refresh patterns in Idempotency and Drift Detection in Python IaC. Define explicit alert thresholds for unauthorized changes.

Configure State Backends & Drift Detection Pipelines Configure State Backends & Drift Detection Pipelines: Configure State then Drift Detection then Drift Detection then Python IaC then IAM Configure State Drift Detection Drift Detection Python IaC IAM
Configure State Backends & Drift Detection Pipelines: the stages run left to right — Configure State, Drift Detection, Drift Detection, Python IaC, IAM.

CLI: pulumi stack select prod CLI: cdktf diff --stack prod CLI: pulumi preview --diff --expect-no-changes

# infra/core/state_manager.py
from __future__ import annotations

import time
import logging
from typing import Protocol, Optional, Callable, TypeVar, Generic
from dataclasses import dataclass
from botocore.exceptions import ClientError

T = TypeVar("T")

class StateBackendProtocol(Protocol):
    def acquire_lock(self, lock_id: str) -> bool: ...
    def release_lock(self, lock_id: str) -> None: ...
    def read_state(self) -> bytes: ...
    def write_state(self, payload: bytes) -> None: ...

@dataclass
class StateOperationResult(Generic[T]):
    success: bool
    data: Optional[T] = None
    error: Optional[str] = None

def execute_with_lock(
    backend: StateBackendProtocol,
    operation: Callable[[], T],
    lock_id: str,
    max_retries: int = 3,
    backoff_factor: float = 2.0,
) -> StateOperationResult[T]:
    """Execute state operations with distributed locking and exponential backoff."""
    for attempt in range(max_retries):
        acquired = False
        try:
            if not backend.acquire_lock(lock_id):
                raise RuntimeError(f"Lock contention on {lock_id}")
            acquired = True
            result = operation()
            return StateOperationResult(success=True, data=result)
        except ClientError as e:
            logging.warning(f"State backend error (attempt {attempt + 1}): {e}")
            if attempt == max_retries - 1:
                return StateOperationResult(success=False, error=str(e))
            time.sleep(backoff_factor ** attempt)
        finally:
            if acquired:
                backend.release_lock(lock_id)
    return StateOperationResult(success=False, error="Max retries exceeded")
  • Verify state file encryption at rest and IAM access boundaries
  • Test concurrent lock acquisition with simulated parallel runs
  • Parse drift output JSON to flag manual console modifications

4. Establish Testing Boundaries & CI/CD Gates

Unit tests must mock provider APIs. Integration tests require isolated sandbox accounts. Gate merges on successful dry-run execution. Pre-commit hooks enforce formatting and static analysis before code reaches the pipeline.

Establish Testing Boundaries & CI/CD Gates Establish Testing Boundaries & CI/CD Gates: moto then localstack then CD Gates then Mock then API moto localstack CD Gates Mock API
Establish Testing Boundaries & CI/CD Gates: the stages run left to right — moto, localstack, CD Gates, Mock, API.

CLI: pytest -m unit tests/ CLI: cdktf deploy --auto-approve --stack staging CLI: pre-commit run --all-files

Separate test environments explicitly. Never run integration tests against production accounts. Mock cloud responses using moto or localstack. Assert zero resource leaks in teardown fixtures.

# infra/tests/test_drift.py
from __future__ import annotations

import time
import pytest
from typing import Protocol, Mapping, Any, TypedDict
from dataclasses import dataclass

class AlertWebhookProtocol(Protocol):
    def send(self, payload: Mapping[str, Any]) -> bool: ...

class DriftPayload(TypedDict):
    resource_id: str
    expected_state: str
    actual_state: str
    severity: str

@dataclass
class DriftDetector:
    webhook: AlertWebhookProtocol

    def parse_diff(self, raw_diff: Mapping[str, Any]) -> list[DriftPayload]:
        """Extract drift events from provider diff output."""
        changes = raw_diff.get("changes", [])
        return [
            DriftPayload(
                resource_id=item["id"],
                expected_state=item["expected"],
                actual_state=item["actual"],
                severity="critical" if item["type"] == "manual_override" else "warning",
            )
            for item in changes
            if item.get("drift_detected", False)
        ]

    def route_alerts(self, drifts: list[DriftPayload]) -> int:
        """Route validated drift payloads to monitoring endpoints."""
        routed = 0
        for drift in drifts:
            sanitized_payload = {
                "resource": drift["resource_id"],
                "severity": drift["severity"],
                "timestamp": int(time.time()),
            }
            if self.webhook.send(sanitized_payload):
                routed += 1
        return routed

class MockWebhook:
    def send(self, payload: Mapping[str, Any]) -> bool:
        assert "pii" not in str(payload).lower()
        return True

@pytest.mark.unit
def test_drift_parser_filters_manual_overrides() -> None:
    synthetic_diff = {
        "changes": [
            {"id": "vpc-123", "expected": "active", "actual": "active", "drift_detected": False},
            {
                "id": "sg-456",
                "expected": "allow_443",
                "actual": "allow_all",
                "drift_detected": True,
                "type": "manual_override",
            },
        ]
    }
    detector = DriftDetector(webhook=MockWebhook())
    results = detector.parse_diff(synthetic_diff)
    assert len(results) == 1
    assert results[0]["severity"] == "critical"
  • Mock cloud API responses with moto or localstack
  • Assert zero resource leaks in pytest teardown fixtures
  • Validate PR checks pass before main merge and block on pulumi preview failures

5. Execute Safe Rollbacks & Production Troubleshooting

State corruption requires surgical intervention. Export state before destructive operations. Verify resource IDs match previous known-good snapshots. Trace provider SDK error codes for retry logic and rate limit handling.

Execute Safe Rollbacks & Production Troubleshooting Execute Safe Rollbacks & Production Troubleshooting: Where it breaks with 4 facets. Where it breaks Execute Safe R watch this boundary Production Tro watch this boundary State watch this boundary SDK watch this boundary
Execute Safe Rollbacks & Production Troubleshooting: the boundaries where things break and what to check.

CLI: pulumi stack history CLI: pulumi stack export > state_backup_$(date +%s).json CLI: pulumi stack import --file state_backup.json

Audit state diffs before executing imports. Implement versioned deployments using commit hashes. Define manual override procedures for emergency bypasses.

Be precise about what a rollback actually is, because two different operations get the same name. Rolling back the code means deploying an earlier commit: the engine computes a new plan and converges the infrastructure back. That is safe and is what you want in almost every case. Rolling back the state means pulumi stack import of an earlier checkpoint, which changes what Pulumi believes exists without touching a single cloud resource. Do that only when the checkpoint itself is wrong — a partially failed update, a corrupted secret, a resource deleted out of band — and never as a substitute for a code revert. Importing a stale checkpoint over a live environment tells the engine that resources created since the snapshot do not exist, and the next update will happily create duplicates of them.

The surgical tools are worth knowing before an incident rather than during one. pulumi state delete <urn> removes a single resource from the checkpoint without deleting it in the cloud — the correct move when a resource was destroyed manually and the engine keeps failing on it. pulumi refresh --target <urn> reconciles one resource rather than the whole stack, which matters when a full refresh takes twenty minutes you do not have. pulumi up --target <urn> --target-dependents applies a fix to one subtree. Every one of those commands narrows the blast radius, and every one should be run only after pulumi stack export has written a snapshot to somewhere outside the repository.

  • Verify rollback restores exact resource IDs and metadata
  • Audit state diff before executing stack import
  • Monitor provider SDK error codes for retry logic and rate limit handling

Operational Notes

Directory layout is the part of "structure" that is easy to see and easy to fix. Stack decomposition is the part that is expensive to change later, because a stack boundary is also a state boundary, a lock boundary, and a permissions boundary. The useful heuristic is rate of change rather than ownership or cloud service.

Stack boundaries follow change frequency Stack boundaries follow change frequency: choose among 3 options. How often does this change? rarely Foundation:accounts, DNS,network weekly Platform: clusters,databases, IAM daily Service: apps andtheir config
Resources that change together belong in one stack; resources that change at different rates do not.

Watch the plan time, because it is your feedback loop. A stack that takes eight minutes to preview will be previewed less often, and a stack nobody previews is a stack that drifts. Once a preview crosses roughly five minutes, split it — the cost of a cross-stack reference is a few lines of code, and the cost of a slow feedback loop is measured in incidents.

Name resources from a single function, not from f-strings scattered across modules. A resource_name(component: str, role: str) -> str helper in core/ that encodes project, environment, and role gives you predictable names, a single place to change the convention, and a cheap way to enforce length limits — several AWS resource types truncate silently at 32 or 64 characters, and a truncated name that collides between two environments is a genuinely nasty failure.

Keep environment differences to data, not to branches. envs/prod.py should differ from envs/dev.py in the values it passes, never in the components it instantiates. The moment production has a code path that staging does not exercise, staging has stopped being a test of production. Where instance sizes or replica counts differ, express them as fields on a validated config model so an invalid combination fails at load time.

Upgrade provider SDKs on a schedule, in their own pull request. A provider bump changes the schema, and the schema changes the plan. Bundling a provider upgrade with a feature change makes the resulting diff unreadable and hides replacements inside noise. One pull request, one provider, with the preview output attached — the dependency-pinning mechanics are in managing IaC dependencies with Poetry and pip-tools.

Make the preview output the review artifact. Reviewers cannot infer a plan from a Python diff, especially where a loop or a conditional changes how many resources are produced. Post the pulumi preview or cdktf diff summary as a pull-request comment and require an explicit acknowledgement for any replacement. That single practice catches more production incidents than any static analysis you can add.

Common Mistakes

Common Mistakes Common Mistakes: Where it breaks with 3 facets. Where it breaks typing watch this boundary mypy watch this boundary IAM watch this boundary
Common Mistakes: the boundaries where things break and what to check.
  • Using global variables for stack configuration, causing cross-environment state pollution.
  • Omitting from __future__ import annotations and relying on runtime typing checks instead of static mypy analysis.
  • Hardcoding provider credentials or backend endpoints instead of environment variables or secret managers.
  • Skipping pulumi stack export backups before destructive updates, leading to unrecoverable state corruption.
  • Running integration tests against production accounts without explicit sandbox isolation and IAM boundary enforcement.

FAQ

How do I prevent state drift when scaling Python IaC across multiple AWS accounts? Implement centralized remote state with DynamoDB locking. Enforce read-only IAM roles for preview stages. Schedule automated pulumi preview --diff --expect-no-changes or cdktf diff jobs with Slack/PagerDuty routing. Parse drift JSON to trigger automated remediation workflows.

What is the safest rollback procedure for a failed Python IaC deployment? Export the last known good state. Verify resource IDs match the target environment. Run pulumi stack import or re-synthesize with CDKTF from the previous commit. Execute a targeted destroy/apply cycle on the affected component only. Never import unverified state.

How do I enforce strict typing in Pulumi/CDKTF components without breaking provider SDK compatibility? Wrap provider inputs in pydantic models or dataclasses with explicit type annotations. Use typing.cast() only when necessary for SDK interoperability. Validate all inputs at stack initialization time. Reject Any or untyped dictionaries in public component signatures.

Should I use Pulumi or CDKTF for large-scale Python infrastructure projects? Choose Pulumi for native Python SDKs, dynamic resource graphs, and direct cloud API integration without a synthesis step. Choose CDKTF for Terraform ecosystem compatibility and when your team has existing Terraform module investments. Both require identical directory structures, state management discipline, and testing boundaries.

How many resources is too many for a single stack? Resource count is a poor proxy; preview time is the number that matters. A stack of 400 small resources may preview in ninety seconds while one of 80 resources with slow reads takes ten minutes. Split when the preview stops being something an engineer will wait for, and split along the change-frequency boundary rather than by cloud service.

Should each team get its own repository for infrastructure? Start with one repository and separate stacks — cross-stack references within a repository are a function call, while across repositories they become a release process. Split into separate repositories only when the ownership boundary is real enough that you would genuinely reject the other team's pull requests, and expect to pay for it in coordination.

How do I share values between stacks without hard-coding them? Export them from the producing stack and read them with a stack reference in the consumer, which keeps the dependency explicit and versioned with the state rather than with the code. Hard-coded ARNs and VPC ids are the single most common cause of an environment that cannot be rebuilt from scratch.