Hybrid Cloud Architecture Cheat Sheet
Covers hybrid cloud connectivity options, common use cases, and key services like AWS Direct Connect and Azure Arc.
Common Hybrid Cloud Use Cases
Why organizations run hybrid architectures.
- Data Sovereignty- Keeping regulated data on-premises while using cloud for compute/burst capacity
- Cloud Bursting- Running baseline load on-prem, scaling into the cloud during demand spikes
- Gradual Migration- Running workloads in both environments during a phased migration
- Disaster Recovery- Using cloud as a DR target for on-prem primary systems
- Latency-Sensitive Edge Workloads- Processing near the data source on-prem, aggregating in the cloud
Connectivity Options
How on-prem and cloud environments are networked together.
- Site-to-Site VPN- Encrypted tunnel over the public internet, quick to set up, variable latency
- AWS Direct Connect / Azure ExpressRoute- Dedicated private network connection, low latency, higher cost and lead time
- AWS Outposts- AWS-managed hardware installed on-prem, running native AWS APIs locally
- Azure Arc- Extends Azure management (policy, monitoring) to on-prem, multi-cloud, and edge resources
- Google Anthos- GKE-based platform for managing Kubernetes workloads across on-prem and multiple clouds
Site-to-Site VPN in Terraform (AWS)
Establishing an IPsec VPN connection between a VPC and on-prem gateway.
resource "aws_customer_gateway" "onprem" { bgp_asn = 65000 ip_address = "203.0.113.10" # on-prem public IP type = "ipsec.1"}resource "aws_vpn_gateway" "vgw" { vpc_id = aws_vpc.main.id}resource "aws_vpn_connection" "main" { customer_gateway_id = aws_customer_gateway.onprem.id vpn_gateway_id = aws_vpn_gateway.vgw.id type = "ipsec.1" static_routes_only = true}
Azure ExpressRoute Circuit + Private Peering (Terraform)
Provisioning a dedicated ExpressRoute circuit and private peering for a hybrid workload requiring predictable low latency.
resource "azurerm_express_route_circuit" "main" { name = "er-circuit-onprem" resource_group_name = azurerm_resource_group.hybrid.name location = azurerm_resource_group.hybrid.location service_provider_name = "Equinix" peering_location = "Silicon Valley" bandwidth_in_mbps = 200 sku { tier = "Standard" family = "MeteredData" }}resource "azurerm_express_route_circuit_peering" "private" { peering_type = "AzurePrivatePeering" express_route_circuit_name = azurerm_express_route_circuit.main.name resource_group_name = azurerm_resource_group.hybrid.name peer_asn = 65001 primary_peer_address_prefix = "192.168.10.0/30" secondary_peer_address_prefix = "192.168.10.4/30" vlan_id = 100}
Hybrid Identity Patterns
Keeping a single source of truth for identity across on-prem and cloud, avoiding the split-brain that breaks SSO.
- Azure AD Connect- Syncs on-prem Active Directory objects and password hashes into Entra ID on a scheduled cycle (default 30 min)
- Pass-Through Authentication- Validates credentials against on-prem AD in real time instead of syncing hashes, keeping password policy enforcement local
- Federation (AD FS / SAML)- On-prem identity provider issues tokens the cloud trusts directly, no credential sync required
- AWS IAM Identity Center + AD Connector- Proxies authentication to an on-prem AD without replicating the directory into AWS
- Conditional Access- Cloud-side policy engine (device compliance, location, risk score) layered on top of federated or synced identity
Connecting an On-Prem Kubernetes Cluster to Azure Arc
Onboarding an on-prem cluster so it appears in Azure Resource Manager for unified policy, monitoring, and GitOps.
az extension add --name connectedk8saz connectedk8s connect \ --name onprem-cluster-01 \ --resource-group hybrid-rg \ --location eastus# Enable Azure Policy and Azure Monitor extensions on the connected clusteraz k8s-extension create \ --name azurepolicy \ --cluster-name onprem-cluster-01 \ --resource-group hybrid-rg \ --cluster-type connectedClusters \ --extension-type Microsoft.PolicyInsights# Deploy via GitOps (Flux) once connectedaz k8s-configuration flux create \ --name platform-config \ --cluster-name onprem-cluster-01 \ --resource-group hybrid-rg \ --cluster-type connectedClusters \ --url https://github.com/org/platform-gitops \ --branch main --kustomization name=infra path=./clusters/onprem
Hybrid-Specific Failure Modes
Problems that don't exist in a single-environment deployment but are common in hybrid architectures.
- Split-Brain DNS- On-prem and cloud resolvers disagree on a name, routing traffic to the wrong side; fix with conditional forwarders (Route 53 Resolver, Azure Private DNS Resolver)
- Asymmetric Routing- Return traffic takes a different path than the request (e.g. out via VPN, back via Direct Connect), tripping stateful firewalls
- MTU Mismatch / Black-Hole Fragmentation- IPsec overhead shrinks effective MTU below 1500; packets silently drop if ICMP fragmentation-needed is blocked
- Clock Drift- On-prem NTP source diverges from cloud time service, breaking Kerberos auth and certificate validation across the link
- BGP Route Flapping- Unstable on-prem router advertisements cause the cloud side to repeatedly recompute paths, adding latency spikes
- Single Circuit Dependency- One Direct Connect/ExpressRoute circuit with no VPN failover means a fiber cut takes down the entire hybrid link
BGP-Based Failover Between Direct Connect and VPN Backup
Router configuration pattern that prefers the dedicated Direct Connect path but fails over to a site-to-site VPN by AS-path prepending.
! On-prem router: Direct Connect BGP session (preferred path)router bgp 65000 neighbor 169.254.1.1 remote-as 64512 neighbor 169.254.1.1 description AWS-DirectConnect network 10.0.0.0 mask 255.255.0.0! VPN BGP session, path prepended to make it less preferredrouter bgp 65000 neighbor 169.254.2.1 remote-as 64512 neighbor 169.254.2.1 description AWS-VPN-Backup neighbor 169.254.2.1 route-map PREPEND-VPN outroute-map PREPEND-VPN permit 10 set as-path prepend 65000 65000 65000! Cloud side automatically prefers the Direct Connect route (shorter AS path)! and fails over to VPN within BGP hold-timer seconds if DX drops
Design hybrid DNS resolution early (e.g. Route 53 Resolver forwarding rules) — most hybrid connectivity failures are actually name resolution failures between on-prem and cloud DNS zones, not routing issues.