• VCF Fleet with Fault Domains and Disaster Recovery Design

    This article is contunuation of the previouse one: VCF Fleet with Disaster Recovery (Across Regions)

    Finally, we bring these concepts together into a combined design that uses both fault domains and disaster recovery.

    Here we:

    • Use availability zones and stretched clusters within a region to provide site high availability and continuous availability.

    • At the same time, we replicate to another region to provide disaster recovery.

    VCF Fleet with Fault Domains and Disaster Recovery Design
  • VCF Fleet with Disaster Recovery (Across Regions)

    This article is contunuation of the previouse one: VCF Fleet with Site High Availability (Across Zones)

    Now let’s move from high availability within a region to disaster recovery across regions.

    On this slide we have:

    • Region 1, which acts as the primary site, and

    • Region 2, which acts as the disaster recovery site.

    VCF Fleet with Disaster Recovery Across Regions
  • VCF Fleet with Site High Availability (Across Zones)

    This article is contunuation of the previouse one: VCF Fleet Basic with Multiple VCF Instances

    Next, let’s move from basic fleet management to site‑level high availability within a region.

    This design uses what you may already know as a stretched cluster:

    • We have a single VCF instance spanning two availability zones.

    • Compute and storage are stretched between the zones, giving us continuous availability and IP portability across them.

    VCF Fleet with Site High Availability Across Zones
  • VCF Fleet Basic with Multiple VCF Instances

    This article is contunuation of the previouse one: VCF Fleet Basic

    Once you’ve deployed that basic VCF instance with Ops and Automation, you can scale out to multiple VCF instances.

    On the icture below you can see several VCF instances, each with:

    • Its own workload domain components – vCenter, NSX, and storage.

    • Its own management domain where required.

    VCF Fleet Basic with Multiple VCF Instances
  • VCF Fleet Basic

    This is first of 5 Fleet design patterns.

    This is core concept of a VCF Fleet Basic deployment.

    Think of this as our baseline VCF instance design – the foundation for all the other VCF Fleet deployment options I’ll cover next.

     VCF Fleet Basic

  • VMware released video for upgrade from vSphere 8.x with LCM to the VCF 9.0

    VMware has released a new video demonstrating the end-to-end upgrade journey from an existing vSphere 8.x environment with Aria Suite to the latest VMware Cloud Foundation (VCF) 9.0 platform.

    LINK TO VIDEO

    The walkthrough showcases an environment that initially relies on Aria Suite Lifecycle Manager and Aria Operations, and illustrates how it is transitioned to the new VCF 9 operating model.
    The process begins with upgrading Aria Lifecycle Manager to version 8.18 Patch 2 or later, which is a mandatory prerequisite. Using this updated lifecycle manager, Aria Operations is then upgraded to VCF Operations 9.0, during which a new component—the VCF 9 Fleet Management Appliance—is automatically deployed.

  • Fault Domains - Availability Zones and Regions

    Standardizing Terminology Across Cloud and VCF

    The key idea is to avoid inventing proprietary terminology and instead rely on concepts that are already widely understood across the cloud industry. Terms such as Region and Availability Zone are common reference points across AWS, Google Cloud, Azure, and other platforms—and they translate well into VMware Cloud Foundation (VCF) architectures.

    By using familiar terminology, we make architecture discussions clearer, more portable, and easier to align with existing cloud operating models.


    Fault Domains

  • VCF 9: Operational Transformation and Platform Convergence

    When we look at VMware Cloud Foundation (VCF) operations today, it’s clear that there have been significant and deliberate changes aimed at simplifying operations while expanding platform capabilities.

    1. Major Improvements in Operations and Troubleshooting

    VCF now delivers a much more streamlined operational and troubleshooting experience, primarily through a centralized management console. Key operational areas have been significantly enhanced:

    • Centralized management console that reduces tool sprawl

    • Improved password and credential management

    • Dramatically enhanced lifecycle management (LCM), including support for VMware ESXi Live Patch, which minimizes downtime and operational risk

    • Centralized license management, addressing challenges introduced by the new Broadcom operating and subscription model

    While some elements—such as registration workflows and licensing portals—are still in their early iterations and continue to evolve, each release brings improved stability, usability, and billing transparency.

    VCF Operations

     

  • VMware Cloud Foundation Architecture poster

    The VMware Cloud Foundation Architecture poster has been comprehensively refreshed to reflect the major innovations introduced with VCF 9, marking a significant evolution of the platform. It delivers a powerful visual narrative that brings the full software-defined data center and cloud operating model to life.

    vmware cloud foundation architecture poster v01 

    Download full version

    More than a diagram, the poster acts as a strategic blueprint for understanding how VCF enables and operates modern private cloud environments. It clearly demonstrates how the platform supports traditional virtual machines alongside cloud-native and next-generation workloads, including Kubernetes, AI, and big data, across multiple data centers and edge locations.

    At its core, the poster showcases how compute, storage, networking, and security are tightly integrated into a single, automated, and lifecycle-managed stack. It illustrates how VCF is deployed, scaled, and operated as a true end-to-end private cloud platform, delivering cloud-like agility with enterprise-grade control.

  • NSX Edge Design in VMware Cloud Foundation

    Why Dedicated vSphere Clusters for NSX Edge VMs Matter

    Executive Summary

    In VMware Cloud Foundation (VCF), NSX Edge Virtual Machines (VMs) play a critical role in delivering north–south and east–west networking services such as Tier-0/Tier-1 routing, NAT, load balancing, and VPN. Because NSX Edge VMs are performance-sensitive and foundational to overall platform availability, their placement and host design are crucial architectural decisions.

    This article is attmpt to explain why dedicated vSphere clusters for NSX Edge VMs are a best practice in VCF, and why collapsing Edge and compute workloads into shared clusters does not provide real cost savings and often introduces operational and availability risks.


    The Cost Myth: Shared Edge and Compute Does Not Reduce Resource Consumption

    A common misconception is that placing NSX Edge VMs on shared compute clusters reduces infrastructure cost by avoiding dedicated hosts. In reality, this approach does not reduce overall CPU or memory consumption.

    Key considerations include:

    • NSX Edge VMs require a fixed amount of CPU and memory to deliver predictable performance. Whether they run on shared or dedicated hosts, the resource demand remains unchanged.

    • Achieving equivalent Edge performance on shared, non-tuned ESXi hosts often requires:

      • Deploying more NSX Edge VMs

      • Spreading them across a larger number of ESXi hosts

      • Limiting placement to no more than one Edge VM per ESXi host to avoid contention

    These requirements quickly negate any perceived cost benefit of collapsing Edge and compute workloads.

    Additionally, performance tuning applied to optimize NSX Edge VMs—such as CPU pinning or enhanced datapath configurations—can negatively impact general-purpose workloads and reduce the achievable consolidation ratio of the cluster.


    Simplification Through Dedicated Edge Clusters

    Deploying dedicated ESXi hosts and clusters for NSX Edge VMs significantly simplifies the ESXi configuration and operational model.

    Benefits include:

    • Tailored host configuration specifically optimized for NSX Edge workloads

    • Ability to adopt the recommended 4-pNIC design, ensuring proper traffic separation and redundancy

    • Enablement of Enhanced Datapath (EDP – interrupt mode) on a controlled subset of hosts without impacting application workloads

    • Reduced configuration complexity compared to maintaining mixed host profiles in a shared cluster

    In contrast, shared environments require compromises that often prevent full adoption of NSX best-practice configurations.


    Improved Performance Troubleshooting and Operational Clarity

    Troubleshooting network performance issues is inherently more complex when NSX Edge VMs share clusters with dynamic application workloads.

    Challenges in shared clusters include:

    • Frequent NSX Edge VM movement due to DRS rebalancing

    • Difficulty correlating performance issues to specific host-level contention

    • Reduced visibility when Edge and compute workloads compete for the same resources

    Dedicated Edge clusters provide:

    • Stable VM placement, simplifying root cause analysis

    • Predictable performance baselines

    • Clear operational boundaries between networking infrastructure and application workloads


    Reduced vMotion Dependency and Improved Availability

    NSX Edge VMs are highly sensitive to vMotion events. Each vMotion operation can temporarily affect packet processing and, by extension, the availability of many dependent workloads.

    With dedicated Edge clusters:

    • NSX Edge VM vMotion is minimized and typically occurs only during planned ESXi maintenance

    • DRS does not need to frequently rebalance Edge VMs due to unrelated workload churn

    • The blast radius of any maintenance activity is reduced and more predictable

    In shared clusters, frequent workload changes often trigger DRS-initiated vMotion events for NSX Edge VMs, increasing the risk of service disruption.


    Additional Optimization Opportunities

    Because NSX Edge VMs have minimal or no dependency on storage services such as vSAN and rarely require vMotion:

    • Network I/O Control (NIOC) can be disabled as a performance-tuning option

    • Host configurations can be streamlined specifically for packet processing efficiency

    • Operational policies can be optimized for infrastructure services rather than application agility

    These optimizations are difficult—or unsafe—to apply in shared clusters hosting diverse workloads.


    Conclusion

    In VMware Cloud Foundation, NSX Edge VMs are infrastructure services, not general-purpose workloads. Treating them as such by deploying dedicated vSphere clusters delivers:

    • No increase in total infrastructure cost

    • Simpler and cleaner ESXi host configurations

    • Improved performance consistency

    • Easier troubleshooting

    • Reduced operational risk and improved availability

    For these reasons, VMware recommends dedicated NSX Edge clusters as the preferred and most robust design for production VCF environments.

  • VCF Architecture - Multi Availability Zone (AZ) Deployment using stretched Clusters

    Modern enterprise workloads demand high availability, resilience, and automated recovery across geographically separated locations. VMware Cloud Foundation (VCF) addresses these requirements through Multi Availability Zone (AZ) deployments using vSphere and vSAN stretched clusters. This architecture enables continuous service availability even during an entire Availability Zone outage, while maintaining centralized management and operational simplicity.

     Multi AZ

    What Is a Stretched Cluster in VCF?

    A vSphere stretched cluster is a single vSphere cluster with vSAN storage that spans multiple Availability Zones (AZs). Each AZ is treated as an independent Failure Domain, typically representing a separate data center, building, or fault-isolated zone with independent power, cooling, and network paths.

    In a stretched cluster architecture:

    • Compute hosts are distributed across two AZs.

    • vSAN synchronously replicates data between AZs.

    • A vSAN witness node provides quorum to prevent split-brain scenarios.

    This design allows workloads to remain available and consistent even if one AZ becomes unavailable.


    Key Architectural Components

    1. Availability Zones as Failure Domains

    Each Availability Zone acts as a vSAN Fault Domain, ensuring that data replicas are placed across AZs. This guarantees that the loss of an entire AZ does not impact data availability or integrity.

    2. vSAN Stretched Storage

    vSAN uses synchronous replication between AZs to maintain identical data copies. This ensures:

    • Zero data loss (RPO = 0)

    • Transparent storage access for virtual machines

    • Consistent performance under normal operating conditions

    3. vSAN Witness Node

    A vSAN witness node is mandatory for stretched clusters. It:

    • Maintains quorum during failures

    • Resides in a third fault location (separate from both AZs)

    • Stores only metadata, not workload data

    The witness ensures deterministic behavior during AZ failures and enables automated recovery.


    High Availability and Automated Recovery

    One of the major benefits of stretched clusters in VCF is the retention of native vSphere features:

    • vSphere HA automatically restarts affected virtual machines in the surviving AZ in the event of an AZ failure.

    • DRS (Distributed Resource Scheduler) continues to balance workloads dynamically based on available compute resources.

    • No manual intervention is required for workload recovery, significantly reducing operational risk and recovery time.

    This results in a fully automated and highly resilient platform capable of surviving complete AZ outages.


    Deployment Order and Best Practices in VCF

    Management Domain First

    In VCF, the Management Domain cluster must be stretched first before stretching any VI Workload Domain (WLD) vSAN clusters. This is critical because:

    • The Management Domain hosts core infrastructure services such as SDDC Manager, vCenter, NSX, and lifecycle management components.

    • Ensuring its availability is foundational for the entire VCF stack.

    Stretch VI Workload Domains as Needed

    After the Management Domain is successfully stretched:

    • Additional vSAN-based VI Workload Domain clusters can be stretched selectively.

    • Not all clusters need to be stretched—only those hosting workloads with strict availability and resiliency requirements.

    • This approach allows organizations to balance cost, complexity, and availability.


    Use Cases and Benefits

    Multi-AZ stretched clusters in VCF are ideal for:

    • Mission-critical applications requiring continuous availability

    • Environments with zero or near-zero downtime requirements

    • Enterprises seeking simplified disaster avoidance rather than traditional disaster recovery

    Key benefits include:

    • Protection against entire AZ failures

    • Zero data loss and rapid recovery

    • Fully automated operations using native vSphere capabilities

    • Consistent management through VMware Cloud Foundation


    Conclusion

    VCF Multi Availability Zone deployments using stretched clusters provide a powerful, resilient architecture for modern private cloud environments. By leveraging vSphere HA, DRS, and vSAN stretched storage—with the correct deployment order starting from the Management Domain—organizations can achieve enterprise-grade availability with minimal operational overhead.

    This architecture transforms Availability Zones into active-active infrastructure components, ensuring business continuity even in the face of major infrastructure failures.

     
     
     
     
     
     

Google AdSence

AUST IT - Computer help out of hours, when you need it most.

Find out why we do it for less.

About

AUST IT will help you resolve any technical support issues you are facing onsite or remotely via remote desktop 24/7. More...

Contacts

Reservoir, Melbourne,
3073, VIC, Australia

Phone: 0422 348 882

This email address is being protected from spambots. You need JavaScript enabled to view it.

Sydney: 0481 837 077

Connect

Join us in social networks to be in touch.

Newsletter

Complete the form below, and we'll send you our emails with all the latest AUST IT news.