Failover Process and Customer Role
Introduction
A failover involves both Superna Support and your own IT teams, each with a distinct role. This article defines what Superna Support's product support entitlement covers during a failover, what your organization is responsible for, how to open and classify a failover support case correctly, and the planning process required to reduce risk ahead of a planned failover.
Roles and Responsibilities During Failover
Superna Support's Role
Superna product support, as it applies to failover, includes:
- The failover planning process, including a readiness health check and remediation ahead of a planned event.
- Failover log analysis during and after the failover, including a post-failover health check and failback assessment.
- Root cause analysis of issues encountered during a failover.
- Documentation of the next steps required to complete a failover under any condition.
What Product Support Does Not Include
- Assisted failover of your data. Hands-on failover services are available separately through Eyeglass Certified partners' professional services.
- Decisions about data protection before, during, or after a failover. These decisions belong to your organization; Superna Support cannot assume responsibility for them.
- Hands-on assistance with data migration using Eyeglass's data and configuration tools. Support can answer questions about use cases, limitations, and recommendations, but will not join a call to perform a migration — this is a data-protection decision that your team must execute. Test all options before attempting a migration.
- Support for third-party hardware and software (Active Directory, DNS, networking, hosts, PowerScale, and similar). Your organization must have its own support agreements and subject-matter expertise for these components. Superna Support can help identify root cause where a third-party component is involved, but is not a substitute for that vendor's support, and cannot join a call or take control of third-party systems for legal reasons.
Your Organization's Role
- Provide (or have access to) all the skills needed to execute and troubleshoot a failover in your environment: Active Directory (including ADSI Edit permissions on computer objects), DNS resolution and updates, PowerScale SyncIQ operations, share/export management, networking and firewalls, Windows sign-in behavior, Linux mount requirements, and application-specific knowledge for anything that depends on NAS shares.
- Stay logged in to the Superna Support portal for the duration of the failover, so you can upload failover logs and communicate with the support team in real time.
How to Open and Classify a Failover Support Case
When you open a support case for a failover, select one of four case types:
| Case Type | When to Use | What It Means |
|---|---|---|
| Not a Failover Case | The case is unrelated to a failover. | Not treated as a failover case. |
| Test Only | Non-production data, no business impact. | Lower priority relative to active production cases. Support assumes the standard failover health-check process, ideally with 7 days' notice; you may opt out if you tell Support when opening the case. |
| Planned | A scheduled, production-data failover for business continuity testing. | Support assumes the failover health-check and remediation process will run ahead of the event, ideally with 7 days' notice. You may opt out, but doing so means accepting the risks the process is designed to eliminate. If you are not on the latest GA release, you are assumed to have read and accepted the risks documented in that release's Release Notes. |
| Unplanned Real DR Event | Production data is affected by an actual, unplanned event. | Highest priority. Only uncheck the Failover Wizard Controlled checkbox if the source cluster shows as unreachable on the LiveOps dashboard. Providing the failover log is mandatory to receive support. This case type must not be used for testing — for testing an uncontrolled failover scenario, use the supported Simulated Disaster Event procedure only; any other test method is not supported. |
Regardless of case type, decisions and execution related to your business data remain your organization's responsibility.
Day-of-Failover Support Process
To get the fastest support response during a failover:
- Update the case when you are about to start the failover so a support engineer can expect the failover log soon.
- Stay logged in to the support portal with the case open, and copy/paste the running failover log from the Failover Wizard into the case as it progresses — do not use email for status updates. If you have questions about anything in the log, paste the log (even if the failover isn't finished) and Support can respond in real time. Starting in release 2.5, the failover log posts a message per policy once that policy's steps are complete and its data access is ready to be tested — start validating data access for that policy as soon as you see the message, and post the log to the case for confirmation.
- Upload the completed failover log immediately once support requests it. This is mandatory to receive support for any question.
- Create and upload a full appliance backup immediately after the failover. Superna Support's automated failover analysis tooling can review hundreds of potential issues from this backup within minutes — much faster than a live call, which is why Support will not join a screen-share in place of reviewing this analysis.
- If a rapid response is warranted, Support may open a chat window (available only while logged in to the support portal) to communicate faster while continuing to review the automated analysis.
Once data access on the target cluster is confirmed and you've posted that confirmation to the case, Support moves on to assess failback steps if any failover steps reported errors.
If you plan to fail back the same day, upload a fresh full appliance backup before failback so Support can assess failback readiness — this is mandatory.
Failover Planning Process (Planned Events)
The failover planning process is a set of steps, split between your team and Superna Support, designed to eliminate known risks before a planned failover.
- Open a support case identifying the event as a planned failover (see case types above). Support will post a failover planning checklist to the case.
- Provide the date, time, and time zone of the planned failover. This lets Support assign a dedicated engineer to actively monitor the case during the event. Without this information, the case is handled under normal support response times instead.
- Allow at least 7 days' notice. This is based on experience across a large volume of prior failovers and gives Support time to address known issues before they affect your event. Opening a case with less notice means accepting that some validations or remediations may not complete in time.
- Submit logs for a readiness health check. The health check covers: Access Zone failover hints; Access Zone, DFS, and Policy readiness status; SPN errors; Eyeglass version and configuration issues; SyncIQ domain mark state; planned failover type; whether the client-redirection guide (DNS or DFS dual delegation) was followed; and Eyeglass appliance and cluster API health. Anything outside this list must be verified independently as part of your own failover plan.
- Complete the acknowledgment items covering your environment (Easy Auditor/Ransomware Defender usage, number of Access Zones or policies failing over, DNS CNAME usage, same-day failback plans, maintenance window length, current RPO reports, and — if applicable — the PowerScale OneFS upgrade procedure) and the one-time acknowledgment items (running the latest supported release or accepting the risk of an older one, reviewing the Failover Release Notes, completing the Failover Planning Guide and Checklist, preparing a contact list, stopping I/O to the source cluster before failover, running domain mark jobs on all SyncIQ policies, and reviewing the operational steps for your chosen failover type and the recovery procedures for failed steps).
This process assumes a scheduled maintenance window: 4 hours of maintenance window is recommended for 25 or fewer SyncIQ policies failing over, and 6 hours for more than 25 policies.
See Also
- Failover Planning – The technical planning checklist referenced throughout the acknowledgment process above.
- Failover Best Practices – DNS, SPN, DFS, and Access Zone design best practices.
- Simulated Disaster Event – The only supported procedure for testing an uncontrolled/unplanned failover scenario.
- Failover Recovery – Recovery steps if a failover does not complete successfully.