Readiness Checks
Introduction
This page has been updated to use 2.15.0 terminology throughout: the Readiness page, with tabs SyncIQ Policy, Access Zones, IP Pools, and Microsoft DFS — see Where things moved for the full legacy-to-2.15.0 mapping. Screenshots on this page and on Failover Validations may still show the legacy DR Dashboard; the underlying readiness concepts and validation criteria are unaffected by the navigation change.
The Readiness page is a vital tool for monitoring and validating the readiness of your disaster recovery environment. It provides a real-time overview of key components such as SyncIQ policies, configuration replication jobs, and network settings. By consolidating critical information in one place, the page allows administrators to quickly assess overall readiness, identify potential issues, and take corrective actions to ensure that all elements are properly configured for failover or failback. It is a display of the last computed result — the status shown was calculated the last time the jobs behind it ran, not a live measurement — and a roll-up: a single status badge stands for a list of individual validations, and the badge alone never tells you what to fix. See Failover Validations for how to read that underlying list.
What is the Readiness page
The Readiness page is a centralized interface designed to monitor and assess the readiness of disaster recovery (DR) systems. It provides real-time validation checks on critical components, such as SyncIQ policies and configuration replication jobs, ensuring that they are properly configured for failover or failback operations. The page categorizes the status of each item into one of five statuses, enabling administrators to quickly identify and resolve potential issues in their DR setup:
- Ready for Failover — every condition the appliance validates has passed. For an Access Zone specifically, this means both required and recommended conditions are met.
- Warning — one or more recommended conditions are unmet. Warning does not block failover — the appliance will let you start one — but the conditions behind a warning can still make that failover fail, so resolve them first rather than reading "not blocked" as "fine."
- Error — one or more required conditions are unmet. Error blocks failover, and the appliance raises a system alarm so the state does not sit unnoticed.
- Disabled — either the configuration replication job is disabled, or the SyncIQ policy itself is disabled in OneFS. Disabled blocks failover. This is worth separating from Error: nothing is broken, something is switched off.
- Failed Over — this item has already been failed over from this cluster, so a failover in the same direction is blocked. Readiness for failing back is not assessed until this failover has completed.
How to Assess Readiness
The Readiness page serves as the main status screen for checking overall cluster readiness. It consolidates key information across all four tabs (SyncIQ Policy, Access Zones, IP Pools, and Microsoft DFS) and sends critical alarms when issues are detected. Below is a guide on how to assess readiness for each tab, as well as a set of general steps that apply to all of them.
Overall Steps for Assessing Readiness
The following high-level steps apply to assessing DR readiness regardless of whether you're dealing with SyncIQ policies, Microsoft DFS, IP Pools, or Access Zones:
-
Log in to the Eyeglass appliance and open the Readiness page.
The Readiness page provides a centralized view of readiness status across all four tabs, with a summary bar naming and counting all five statuses — read the summary bar rather than the table when you want the whole picture, since it names every status regardless of which ones happen to be present that day.
-
Check the Readiness Tabs
Depending on your needs, review the specific tab — SyncIQ Policy, Access Zones, IP Pools, or Microsoft DFS — to evaluate failover preparedness.
-
Verify Overall DR Status
The status badge on each row gives you an immediate view of whether the item is ready for failover. If an item reads Error, Eyeglass has already triggered a system alarm, so you can take corrective action before a failover is attempted.
-
Review Detailed Status
Select the row to open a slide-out listing its Errors and Warnings. From there, select Details to open the item's own Topology & Mapping page for detailed information about any issues or actions required to resolve them — see Failover Validations for how to read that list.
-
Fix Critical Issues
Investigate any Error, Disabled, or Warning states and resolve them to ensure that the system is fully prepared for DR events.
Assessing Microsoft DFS Readiness
Go to the Readiness page and select the Microsoft DFS tab. The status is Ready for Failover when:
- Your SyncIQ Policy is enabled.
- The Last Started and Last Success timestamps of your SyncIQ Policy are identical.
- The Eyeglass configuration replication job is enabled.
- The Last Run and Last Success timestamps of the Eyeglass configuration replication job are the same.
- The audit status of the Eyeglass configuration replication job is OK.
The status is displayed per policy. It also indicates which pair of clusters are used in the DFS configuration and the associated SyncIQ policy.
Expand a policy to see the details for the SyncIQ Policy and the Eyeglass Configuration Replication status:
- Last Run time of the SyncIQ policy and the status of the last run.
- Last Run time of the Eyeglass Configuration Replication job and the status of the last run and audit.
If new shares are created on the DFS-type policy, the next run of the Eyeglass Configuration Replication job in DFS type will be aware of the new shares and ready to fail them over. It is important to check this tab after creating more DFS shares under a policy.
Assessing SyncIQ Policy Readiness
Go to the Readiness page and select the SyncIQ Policy tab to view a summary of the SyncIQ Policy statuses. The status for each policy combines the SyncIQ PowerScale OneFS data replication status and the Eyeglass Configuration Replication job status, providing an overview of the readiness for failover.
This combined status is updated every time the Eyeglass Configuration Replication task runs, which makes it the most reliable single indicator of DR readiness for failover. Administrators can use this information to review the status of each component, identify errors, and resolve them to ensure that all SyncIQ policies are properly configured for failover.
- Ready for Failover means all SyncIQ Policy Failover Recommendations have passed validation, indicating the policy is safe to fail over.
- Error means one or more SyncIQ Policy Failover Recommendations have failed, and the policy is not ready to fail over. Eyeglass triggers a system alarm to notify you of the issue.
- Disabled means the configuration replication job or the SyncIQ policy itself is disabled in OneFS.
If you make any changes to your environment, the following Eyeglass jobs must run before the SyncIQ Policy Readiness status is updated:
- Configuration Replication
Assessing Access Zone Readiness
Go to the Readiness page and select the Access Zones tab to view a summary of key networking, SmartConnect, and Kerberos Service Principal Name (SPN) configurations. The status for each is combined to provide an overall status, checked in both directions of a replicating cluster pair so you get failover and failback status together.
- Ready for Failover means all requirements are met, and the Access Zone is ready for failover.
- Warning means all required conditions are met, but some recommended conditions are not. In this state, the Access Zone can be failed over, but there may be some additional manual steps — the Failover Wizard will allow you to start a failover in this state.
- Error means unmet required conditions, and the zone is not ready to fail over. A system alarm is triggered to block any failover attempts — the Failover Wizard will block you from starting the failover in this state.
If you make a change to your environment, the following Eyeglass jobs must run before the Access Zone Readiness will be updated:
- Configuration Replication
- Failover Readiness
The Failover Readiness job is created automatically between each replicating cluster pair, one job per direction, named <cluster name>_<cluster name> — and its initial state is disabled. If Access Zone readiness has never populated, that is the first thing to check; enable it with igls admin schedules set --id Readiness --enabled true. Once enabled, this job runs on a default interval of 15 minutes.
Assessing IP Pool Readiness
Go to the Readiness page and select the IP Pools tab, covering readiness for IP Pool failover. It updates on the same Configuration Replication + Failover Readiness dependency as Access Zone readiness, described above.
Scope is the thing to remember here: an IP Pool failover is limited to a single Access Zone, while an Access Zone failover can cover several pools at once. That is why the Access Zones and IP Pools tabs can disagree about the same underlying network configuration without either being wrong.
Readiness is not assessed for an Access Zone that has already failed over. The Readiness page only displays the readiness status for failover from the current active (primary) cluster to the DR target cluster. The readiness status for "failback" (returning operations to the original primary cluster) is not checked until after the failover to the target cluster has been completed.
About Validations
Failover Validations are a core component of ensuring your system is ready for disaster recovery. These checks evaluate the status of key components like SyncIQ policies, Access Zones, and network configurations. The results are consolidated in the DR Dashboard to provide a clear picture of whether your environment is safe to failover.
Failover Validations are performed regularly and include checks for data replication, network readiness, configuration consistency, and more. Each validation either confirms that the system is ready or flags critical issues that need to be addressed before failover.
For more details on specific validations, refer to the Failover Validation Reference.
See Also
- Runbook Robot: Automate daily failover and failback testing for continuous, hands-off readiness validation.
- Failover Estimate: Predict failover duration from historical data, surfaced alongside readiness status in the DR Dashboard.