Skip to main content
Migration Notice
We're migrating documentation from the old portal into this one. Some things may look a little different or out of place in the meantime — we know, and we're working to get it right. If something's unclear or doesn't look right, let us know.
Version: 2.15.0

Advanced License and Pipeline Features for Data Orchestration for Dell

Introduction​

This guide covers the Data Orchestration (Golden Copy) features that require the Advanced license, the Backup Bundle license, or the Pipeline license subscription. The features are grouped by license requirement. For general folder, job, and schedule configuration, see the Configuration Guide. For self-service archiving and policy-based archiving, see the Archive Engine Guide.

The Advanced license features make backup workflows easier to monitor and automate, and add data integrity features, API access for automation with external tools, and new reporting options.

Requirements:

  • A Golden Copy Advanced license, or a Golden Copy Backup Bundle license, applied to the appliance.
  • Features marked as requiring the Pipeline license subscription need that license in addition.

Assign an Advanced License to a Cluster​

An Advanced license enables the features in this guide for a cluster managed by Golden Copy:

searchctl isilons license --name <clusterName> --applications GCA
note

An available Advanced license must exist for this command to succeed, and the cluster also requires a base Golden Copy license before an Advanced license key can be applied.

PowerScale Cluster Configuration Backup​

Backing up data does not fully protect the cluster configuration that is needed to recover a cluster. This feature collects a backup of the shares, exports, and quotas using the OneFS 9.2 export feature, and stores a copy of the configuration in an S3 bucket configured to store the backup. Configuration backup jobs can be scheduled on the cluster, and incremental change list aware backup picks up the new backup job files and copies them to the bucket.

Requirements: OneFS 9.2 or later.

Configuration backed up:

  • Quotas
  • SMB shares
  • NFS exports
  • S3 configuration
  • Snapshot settings
  • NDMP configuration settings
  • HTTP settings

Run an Immediate Cluster Backup​

  1. Show the command help:

    searchctl isilon config
  2. Back up the configuration immediately. The command returns the task ID of the backup job, which is referenced in the folder name copied to the S3 bucket:

    searchctl isilon config --name <isilon-name> --backup

The EyeglassAdminSR account for Golden Copy must also include the import/export privilege to run this command:

isi auth roles modify EyeglassAdminSR --add-priv ISI_PRIV_CONFIGURATION
note

This only generates a backup on the cluster. It does not copy the backup to S3 until a folder configuration is added and run, as described in the next section.

Enable a Folder to Store a Cluster Backup​

Use the --backup-isilon flag of searchctl archivedfolders add to specify the S3 bucket that stores the cluster backup. This step is required to protect the cluster configuration. The flag is used instead of the --folder path and cannot be set together with --folder. It sets the folder path to ISILON_BACKUP_CONFIGURATIONS_PATH automatically. Complete the other mandatory flags needed to copy data to the S3 bucket.

The following example for AWS runs a daily backup:

searchctl archivedfolders add --isilon <cluster name> --backup-isilon --accesskey xxxxxxx --secretkey yyyyyy --endpoint <region endpoint> --region <region> --bucket <backup bucket> --cloudtype aws --incremental-schedule "5 0 * * *"
note

This backs up the cluster configuration 5 minutes after midnight, which provides time for the cluster backup to be created. In the incremental schedule 5 0 * * *, the first number is the minutes after the hour. The value can be changed if needed.

Schedule the Cluster Configuration Backup​

Create a new backup of the cluster configuration once per day as a best practice. Use NEVER or DISABLED to deactivate the schedule:

searchctl isilon config --name <isilon-name> --set-backup-cron "0 0 * * *"

Restore the Configuration Backup to the Source Cluster​

warning

Only use this command when you need to restore the configuration. It overwrites all shares, exports, and quotas on the named cluster.

  1. Browse the S3 bucket with the Cloud Browser GUI and locate the bucket/clustername/ifs/data/Isilon_Support/config_mgr/backup path. Find the subfolder name that contains the backup you want to use for the restore.

  2. Restore the configuration named by the ID. The command returns the task ID of the restore job, and the job reads the backup data from the bucket and restores it to the source cluster:

    searchctl isilon config --name <isilon-name> --restore <id>

    Example:

    searchctl isilons config --name gcsource93 --restore gcsource93-20220312020124
  3. Optionally, change the path that stores the backup in the S3 bucket. Edit /opt/superna/eca/eca-env-common.conf and add the following variable. The default path is /ifs/data/Isilon_Support/config_mgr/backup. The path is fixed by PowerScale, and the variable exists in case the path changes in the future:

    export ISILON_BACKUP_CONFIGURATIONS_PATH=xxxx

Recall by Timestamp​

Recall can scan the metadata in the object properties to select files based on the created or modified time stamps the files had on the file system when they were backed up.

Limitations: On a folder with an active archive job, recall jobs are blocked until the archive job completes. This prevents recalling data that is not yet copied.

The recall command accepts the following options. Use double quotes, and enter the times in the UTC time zone:

  • --start-time "<date and time>" in the format yyyy-mm-dd HH:MM:SS, for example 2020-09-14 14:01:00.
  • --end-time "<date and time>" in the same format.
  • --timestamps-type {modified, created} selects which time stamp is evaluated. The default is the modified time stamp.

Example that recalls files with a last modified time stamp between September 14, 2020 and September 30, 2020 under the /ifs/xxx folder:

searchctl archivedfolders recall --id 3fd5f459aab4f84e --subdir /ifs/xxx --start-time "2020-09-14 14:01:00" --end-time "2020-09-30 14:01:00" --apply-metadata

Example that selects files by created time stamp:

searchctl archivedfolders recall --id 3fd5f459aab4f84e --subdir /ifs/xxx --start-time "2020-09-14 14:01:00" --end-time "2020-09-30 14:01:00" --timestamps-type created --apply-metadata

Power User Restore Portal​

The Power User Restore portal targets power users who need to log in and recall data but who should not be administrators of the Golden Copy appliance. The portal is aware of SMB share security. After an Active Directory login, the user's SMB share paths are compared to the backed up data, and the visibility of data in the Cloud Browser is restricted accordingly. The user also sees only the icons related to Cloud Browser.

Requirements:

  • Golden Copy node 1 must have SMB TCP port 445 access to the clusters added to the appliance.
  • SMB shares must exist on a path at or below the folder definitions configured in Golden Copy.
  • Admin-only mode must be disabled on the Golden Copy appliance.

User login and recall experience:

  1. The user logs in with user@domain and password.
  2. The options in the UI are limited to Cloud Browser and stats.
  3. The user selects the cloud type and bucket name. The data displayed is based on the user's SMB share permissions.
  4. Clicking the icon launches a recall job.
note

The recall data path is not the original data location. The power user needs NFS or SMB access to the /ifs/goldencopy/recall/ifs/xxxxx path to retrieve the data. Designing a user restore solution requires some planning with SMB shares to grant access to data that has been recalled.

Redirect Recall Data to a Different Cluster​

There are two redirect use cases:

  • Recall to a target cluster with metadata, which requires a license on the target cluster.
  • Recall to a target cluster without metadata, which does not require a license on the target cluster.

Redirect a Recall with Metadata​

  1. Add the target cluster to Golden Copy, put all NFS mounts in place by following the Installation Guide, and apply a valid license to the cluster.

  2. Add a folder definition that specifies Cluster B and the same S3 bucket that is used to back up Cluster A.

  3. Recall data from Cluster A to Cluster B:

    searchctl archivedfolders recall --id xxxx --source-path clusterA/ifs/data/yyyy

    xxxx is the folder ID added in the previous step, and yyyy is the full path of the data that you want to recall to Cluster B but that was backed up from Cluster A.

Redirect a Recall without Metadata​

warning

No metadata means that the owner, group, mode bits, and folder ACLs are not restored.

  1. Create an NFS export with root client mount permissions, on a location on any device, with the IP addresses of each Golden Copy VM. The only supported method to recall data plus file metadata (owner, group, mode bits, folder ACLs) requires PowerScale as the target. Other storage devices require metadata recall to be disabled.

  2. On each Golden Copy node, run sudo -s, edit /etc/fstab to change the device IP address and path to the NFS export created in the previous step, and save the file. Repeat on every Golden Copy node.

  3. For a PowerScale target, start the recall job and specify the source path where the data is stored in the S3 bucket. YY is the new folder ID created when the folder was added. The recall NFS mount must be created on the new target cluster on the /ifs/goldencopy/recall path before a recall can be started:

    searchctl archivedfolders recall --id YY --apply-metadata
  4. For other NFS storage:

    searchctl archivedfolders recall --id YY

The recall job recalls the data to the new cluster added to the folder definition.

Target Object Store Stats Job​

The s3stat job type scans all of the objects protected by a folder definition and reports the object count, the total data stored, the oldest file found, the newest file found, and the age of the objects.

searchctl archivedfolders s3stat --id <folderID>

View the results with the job ID:

searchctl jobs view --follow --id <job ID>
note

Cloud provider storage usage charges are assessed based on the API list and get requests made to the objects. The job can be canceled at any time with the searchctl jobs cancel command.

Data Integrity Job​

The data integrity job audits folders by using the metadata checksum custom property. Random files are selected for audit on the target device, downloaded, and a checksum is computed and compared to the checksum stored in the metadata. If any files fail the audit, the job report summarizes the failures and the successful audits. This verifies that the target device stored the data correctly.

Requirements:

  • Release 1.1.6 update 2 or later.
  • Data must be copied with the global --checksum flag set to ON. If the data did not have a checksum applied during the copy, the job returns a 100% error rate. See the global configuration settings with searchctl archivedfolders getConfig and searchctl archivedfolders configure -h.

Run the data integrity job. x is the total amount of data in GB that is randomly audited, and yy is the ID of the folder to audit:

searchctl archivedfolders audit --folderid yy --size x

Monitor the job with the job ID:

searchctl jobs view --id xxxxx

Media Mode​

Media workflows and applications generate symlinks and hardlink files in the file system to optimize disk space. Backing up these files to objects normally treats each symlink or hardlink as a full-sized unique file, which causes the data set in the object store to grow significantly, and objects recalled back to a file system lose the symlink and hardlink references.

Media mode automatically detects symlinks and hardlinks and backs up the real file only once. It maintains the references to this file to optimize backup performance and to preserve the links in the object store. When the data is recalled, the symlinks and hardlinks are recreated in the file system to maintain the optimized data structure. The mode can be enabled or disabled, and it supports full backup and incremental detection of new symlinks or hardlinks that are created between incremental sync jobs.

Requirements: Pipeline license subscription.

note

The source and target of a hardlink or symlink must be under the archive folder definition.

  1. Log in to Golden Copy node 1 and edit the configuration file:

    nano /opt/superna/eca/eca-env-common.conf
  2. Add the following variables:

    export IGNORE_SYMLINK=false
    export IGNORE_HARDLINK=false
    export ARCHIVE_LINK_SUFFIX=.link.json
    export INDEX_WORKER_RECALL_SKIP_EXISTENT_FILES=true

    INDEX_WORKER_RECALL_SKIP_EXISTENT_FILES=true skips files in the S3 bucket if they already exist in the recall path.

  3. Optionally, add the following variables:

    • Requires release 1.1.20 or later. Skips files that already exist on disk, which allows archive worker restarts during jobs to skip over data that is already downloaded:

      export ARCHIVE_WORKER_RECALL_SKIP_EXISTENT_FILES=true
    • Requires release 1.1.19 or later:

      export IGNORE_ADS_FILES=true
  4. Save and exit the file, then restart the cluster:

    ecactl cluster down
    ecactl cluster up

When media mode is enabled, the job view summary provides the symlink and hardlink count and the MB copied, along with the skipped files that did not need to be copied because they were redundant copies. A calculation of the data copy savings when duplicate data is not copied is also provided. The stats view summary includes detailed stats on inodes, symlinks, hardlinks, and skipped copies when data reduction has occurred because duplicate data was removed from the copy job.

If ARCHIVE_ENABLE_INODE_SKIP_EXISTS is set to true, Golden Copy does not check for the presence of duplicate inodes when it processes hardlinked data. This results in higher bandwidth from Golden Copy to cloud targets. The default is false, and it should only be set to true if directly instructed by Support.

Media Workflows Dropbox Folder or S3 Bucket​

The dropbox feature monitors a folder on the cluster for new files and syncs them to an S3 bucket. The source files are deleted after the copy of the files completes. This enables a workflow in which completed work is dropped into a folder and moved to the production finished-work storage bucket. This feature manages the on-premises to cloud direction. With the Pipeline license, an S3 bucket can also be offered as a dropbox bucket, where new data copied into the bucket is copied to on-premises storage and then deleted from the bucket.

Requirements: Pipeline license subscription, for both the on-premises to cloud and the cloud to on-premises dropbox use cases.

Support limitation: The dropbox delete at source feature is a file-by-file delete designed for media workflows. It is not a high-speed delete for terabytes of data or millions of files.

Behavior to be aware of:

  • Data copied into a dropbox folder path may have many levels of folder structure. The delete process uses one thread per copied file, and the thread deletes the file only if the target storage returns an HTTP 200 return code, which means it is OK to delete the file. The delete occurs in the same thread that copied the file.
  • At the end of the job, many empty folders remain. Golden Copy has a cleanup task to delete these folders, and each folder detected during the job has at least one delete API request executed. If the folder is empty, it is deleted successfully, and the stats view shows the successful deletes of dropbox folders. If the folder is not empty because a file in the folder failed to copy, or because new data was copied into the folder after the job started, the folder delete request fails and the folder is left alone. The failed file can be retried on the next dropbox job.

On-Premises Dropbox​

Apply the delete-from-source rule to every full archive of a folder when the folder is added:

searchctl archivedfolders add <add folder parameters> --delete-from-source

Apply the rule to a single full archive only:

searchctl archivedfolders archive --id <folderid> --delete-from-source

Cloud Bucket Dropbox​

This feature allows a bucket in the cloud to receive data from various sources and have it automatically synced to the on-premises cluster on a schedule. When the sync completes, the source files in the bucket are moved to a trash bucket that has a time-to-live (TTL) lifecycle policy applied. The flag with a bucket name stores the data deleted from the main bucket, so a bucket policy can specify how long deleted data remains before it is purged, which provides a time period to recover data from the trash bucket if needed.

searchctl archivedfolders add --folder /bucket1 --isilon <name of cluster> --source-path '/bucket path' [--pipeline-data-label xxxxx] --recall-schedule "*/30 * * * *" --cloudtype {aws, azure, ecs, other} --bucket <bucket name> --secretkey <> --accesskey xxxxx --trash-after-recall True --recyclebucket <trashbucket>

Long Path Support​

Many media workflows create long path names, which are supported in NFS and POSIX file systems. These path lengths can exceed the S3 object key limit of 1024 bytes, which blocks copying folder data at deep levels in the file system.

The long path feature supports the long path names that exist in the file system and need to be copied to S3. It creates object symlinks for all data that would fail to copy to the S3 bucket.

Requirements:

  • OneFS 9.3 or later, which supports REST API calls for a path that is 4K bytes long.
  • Golden Copy 1.1.7 build 22067 or later.

To enable the feature:

  1. Edit the configuration file:

    nano /opt/superna/eca/eca-env-common.conf
  2. Add the following variable, then save and exit the file:

    export ENABLE_LONGPATH_LINKING=true
  3. Restart the cluster:

    ecactl cluster down
    ecactl cluster up

Pipeline Workflow License​

The Pipeline license enables S3 to file workflows and file to object workflows. The S3 to file direction requires the incremental detection feature, which uses the date stamps on the file system to match the object time stamp. This allows incremental sync from S3 to file, which tracks differences between S3 and the file system and syncs only new or modified objects from the cloud down to the file system.

Typical use cases:

  • Media workflows that pick up media contributions from a third party from a cloud S3 bucket and transfer the data to an on-premises PowerScale for editing workflows.
  • Media workflows that download S3 output from a rendering farm, when the output is needed on-premises for video editing workflows.
  • HPC cloud analysis for AI/ML that requires on-premises data to be copied to an S3 cloud bucket as analysis input, which produces output in a different bucket that needs to be copied back on-premises.
  • Many to one: many source buckets drop data into a common project folder on PowerScale, so data from multiple locations is accepted and dropped into one project folder on the file system.

These workflows support scheduled copies or incremental sync in both directions, between different sources and destinations, to pick up only new or modified files.

Requirements: Apply the Pipeline license key to Golden Copy and assign it to a cluster:

searchctl isilons license --name <clusterName> --applications GCP

Prerequisite Configuration​

Data that is pipelined from the cloud needs to land on the cluster in a path. This requires an NFS mount to /opt/superna/mnt/recallsourcepath/GUID/clusternamehere/, mounted on PowerScale onto a path on the cluster, for example /ifs/fromcloud. This path stores all data arriving from the cloud S3 buckets. The data from the S3 buckets can be shared using SMB or NFS protocols. For the NFS export and mount steps, see the Installation Guide.

Create a Pipeline​

A pipeline is a cloud to on-premises storage configuration. One-to-one or many-to-one definitions are supported to map multiple source S3 buckets to a single folder path target.

  1. Modify the PowerScale cluster that is the target of the cloud pipeline. This setting is attached to the cluster configuration globally for all pipelines that drop data on the cluster. Repeat for each cluster that has pipelines configured. xxxx is the cluster name, and AbsolutepathforFromcloudData is the path used for the from-cloud NFS mount. The default /ifs/fromcloud should be used:

    searchctl isilons modify --name xxxx --recall-sourcepath AbsolutepathforFromcloudData

    Example:

    searchctl isilons modify --name cluster1 --recall-sourcepath /ifs/fromcloud
  2. Create a pipeline configuration from an S3 bucket and path to a file system path on the cluster. This example scans the bucket every 30 minutes and copies only new files to the cluster:

    searchctl archivedfolders add --folder /ifs/fromcloud/bucket1 --isilon <name of cluster> --source-path '/bucket path' --recall-schedule "*/30 * * * *" [--pipeline-data-label xxxxx] --cloudtype {aws, azure, ecs, other} --bucket <bucket name> --secretkey <yyyy> --accesskey <xxxx> --recall-from-sourcepath
    • --folder is the fully qualified path where the data is copied to on the cluster. The path must already exist.
    • --source-path is the S3 path, or bucket prefix, to read data from in the bucket. Use / if the entire bucket is copied.
    • --recall-from-sourcepath enables the Pipeline feature and is mandatory.
    • --recall-schedule is the schedule to scan the S3 bucket for new or modified data.
    • --trash-after-recall <trashbucket name> enables the dropbox feature, which moves the bucket data to the specified trash bucket after the sync from cloud to on-premises completes. The trash bucket should be in the same region as the source bucket.
  3. Run the job to scan the bucket and copy the data from the cloud to the cluster file system:

    searchctl archivedfolders recall --id <folderID>

Fan-In Topology​

Fan-in allows multiple source S3 buckets to be synced to a single path on the on-premises PowerScale cluster. The prefix path must be unique to avoid file system collisions. For example, to sync bucket1 and bucket2:

searchctl archivedfolders add --folder /ifs/fromcloud/bucket1 --isilon <name of cluster> --source-path '/bucket path' --recall-schedule "*/30 * * * *" [--pipeline-data-label xxxxx] --cloudtype {aws, azure, ecs, other} --bucket bucket1 --secretkey <yyyy> --accesskey <xxxx> --recall-from-sourcepath
searchctl archivedfolders add --folder /ifs/fromcloud/bucket2 --isilon <name of cluster> --source-path '/bucket path' --recall-schedule "*/30 * * * *" [--pipeline-data-label xxxxx] --cloudtype {aws, azure, ecs, other} --bucket bucket2 --secretkey <yyyy> --accesskey <xxxx> --recall-from-sourcepath

Job Transaction Log Backup​

This feature stores the success, skipped, and errored details of each job in an S3 bucket to provide an offsite audit log of all actions taken during copy jobs. The log records each action for every file copied and is large.

How it works:

  • When enabled, the archivereporter container monitors all archivereport-<folder-id> topics. When jobs start, newly created messages in the topics are written to temporary text files at time intervals, using a timer. The host path for the temporary logs is /opt/data/superna/var/transactions/.
  • The default time interval is 10 minutes. After every interval, new messages with information about the temporary text files are published to the transactionLog topic.
  • The archiveworker containers spawn a consumer to monitor the transactionLog topic. Messages are picked up individually and sent to the configured S3 logging bucket, and the temporary file is deleted once the upload succeeds. Only the archive worker running on the same container as the archive reporter continues to consume the message.
  • Exported log files are sent to the logging bucket under <folder-id>/<job-type>/<job-id>/<file-name>, with the file name containing the time stamp. Metadata logs are sent under <folder-id>/METADATA/<job-id>/<file-name>.
  • The feature uses the cloud credentials of each added archived folder. Each of those cloud credentials needs a bucket with the name defined by LOGGING_BUCKET.

Configuration: Transaction logging is disabled by default.

  1. Edit the configuration file:

    nano /opt/superna/eca/eca-env-common.conf
  2. Add the following variables. TRANSACTION_WRITE_INTERVAL is the interval in minutes for sending temporary files to the S3 bucket. LOGGING_BUCKET is the S3 bucket that stores the audit transaction logs, and the bucket must be defined in a folder configuration:

    export BACKUP_TRANSACTIONS=true
    export TRANSACTION_WRITE_INTERVAL=10
    export LOGGING_BUCKET='gcandylin-log'
  3. Save and exit the file, then restart the archive reporter service:

    ecactl cluster services restart archivereporter
note
  • The logs can be exported to JSON with the searchctl archivedfolders export command.
  • The archive workers upload by consuming events and uploading directly to the cloud, which consumes bandwidth and storage costs in the target bucket.

Metadata Content Aware Backup​

This feature is content aware, with support for over 1500 file types. When enabled on a per-folder basis, metadata is extracted from the file and added as custom tags to each object, and added to a searchable full content index. This preserves key metadata in the object store, where it is human readable with S3 browser tools and the Cloud Browser built into Golden Copy. It also allows administrators or users to search for data to recall based on the metadata within the files. The additional overhead to extract the metadata is minimal, because it is all done in RAM during the backup process.

For example, a Word document has rich metadata that can be populated. With content aware backup mode, the file attributes are extracted and placed into custom user-defined fields on the object. The GC_xxxx fields represent a dynamic tag extracted from each file that has metadata.

Requirements:

  • This is an experimental feature. Use it with caution.
  • Pipeline license subscription. Run searchctl license list; the Pipeline license must be listed.
  • Release 1.1.11 or later.

Limitations: The flag is supported only with Golden Copy or Data Orchestration edition. It is not supported with Archive Engine workflows.

Configuration: Use the --customMetaData flag when adding or modifying a folder. When data is archived, the metadata from the files is extracted, copied into custom S3 properties, and indexed so it can be searched in the Cloud Browser.

See Also​