Disaster Recovery Procedure

This Disaster Recovery Procedure provides a structured approach that should be followed by the University to identify and assess IT DR incidents, and commence recovery of the critical IT systems, applications and infrastructure in case of a disaster to ensure continual operations and resume mission critical functions.

The safety of all individuals and the protection of property are of paramount importance to the University. Recovery operations are only commenced once any ongoing risk to personnel and property has been minimised/mitigated

The following steps should be used when recovering The University’s IT systems, applications and infrastructure and should be executed in sequence to effectively recover critical business operations.

Notification

First line responders (i.e. IT Operations Service Desk, SES Service Desk) will receive calls, emails or walk-ins from staff or students about a suspected disaster. Alternatively, the infrastructure team will receive system alerts on any issues related to a possible disaster via network monitoring software, email or manual processes.

The IT Operations Service Desk must log and categorise an incident in Systems Center Service Manager (SCSM) and document the information received as described in IT Incident Management Process.

Assessment

IT Operations service desk assess whether identical issues have been reported by others and associate related incidents to a parent incident where necessary. IT Operations service desk assigns severity to the incident / parent incident based on this assessment and performs initial diagnosis of the incident.

If the assigned severity is 1, the Service Desk must notify the CIO and the IT Leadership Group. With the guidance of the CIO, the IT Operations service desk will assign the incident / IT disaster to the Infrastructure team and attach supporting documentation related to the incident.

Recovery

This phase will commence when the CIO, with the support of the IT Leadership Group, activates the IT DRP.

The following steps are to be used when recovering The University’s IT systems, applications and infrastructure at the original or an alternate site. Steps are grouped according to team. Each step should be executed in sequence to maintain efficient operations.

CIO / IT Leadership Team

CIO and the IT Leadership Team have overall accountability for the control and coordination of the disaster recovery effort. They shall determine the appropriate response and required resources based on the incident assessment. This team shall liaise with other business units, external stakeholders and third-party organisations when they are impacted by the disaster to:

  • Recommend initiation of disaster recovery activities
  • Engaged the Infrastructure team with appropriate personnel
  • Determine whether a Communications and Media team is necessary
  • Coordinate disaster recovery activities across all divisions and business units
  • Arbitrate if conflicting recovery objectives arise
  • Approve close-down of recovery environments (Broadway DR site) following a return to normal operations
  • Initiate a post-disaster review and recommend appropriate changes to Disaster Recovery Plans.

Infrastructure Team

The infrastructure team shall have primary responsibility for ensuring appropriate hardware, software and network facilities are available in the recovery process. The team shall ensure that a minimum operating environment is available. Their responsibilities include:

  • Determining what, if any, recovery hardware is required
  • If hardware is required, determining if suitable components are available either from storage, spare capacity, loan equipment from system vendors, etc
  • Depending on which critical systems are impacted, Infrastructure Team Manager contacts appropriate System Vendors to assist / perform recovery effort of the critical systems. For example, if the Learning Management System provided by Backboard is impacted, Blackboard should be contacted to assist in the recovery process as Blackboard is SaaS service and SLAs are in place. Infrastructure team shall refer to the Vendor contact list for the relevant information
  • If necessary, procurement activity shall be commenced to obtain permanent replacement for the failed / destroyed hardware
  • Perform testing over the procured hardware and install the appropriate version of the Operating System and patches, if necessary. If using an existing machine, verify the correct Operating System and necessary patches are installed
  • (Re-)Configure the hardware.  Wherever possible the IP address, server name, SID etc will be configured to match the failed hardware. This will minimise the re-configuration of applications and user machines
  • The Infrastructure team shall obtain the necessary backup copies of the application data from most suitable backup source, either the primary backup system in Fremantle or the DR site in Broadway. If the backup is the sole source for recovery of the system and/or data, a copy will need to be made before the backup is used for any recovery processing
  • Initial configuration and testing of the restored application(s) will be undertaken by Infrastructure team in conjunction with the Resolver groups such as BST and CSAT as required
  • Once the recovery environment has been successfully tested, the Infrastructure team shall inform the CIO and IT Leadership Group of the availability of the application.  This shall include confirming the Recovery Point achieved and application restrictions, if any.

IT Disaster Recovery Plan Testing

Testing the IT DRP ensures that the correct data is being backed up, that the data is restorable and that The University is aware of how to restore it. The University will perform IT DRP testing within a three-year cycle covering all critical systems. A schedule of The University’s IT DRP testing cycle and approach can be found on the website

Per ISACA best practices, The University will use the following principles to prepare for consideration when performing IT DRP testing:

  1. Identify and rank critical applications

    The University will review and identify the critical applications significant to The University’s operations, and consider interfaces, data flows and dependencies.

  2. Create a recovery team with roles and responsibilities

    The University CIO will create a recovery team encompassing all the functions and roles essential to restore the critical applications and IT services quickly and completely. The University will create a document that identifies the team members, their respective roles and the steps each personnel would take in restoring operations.

  3. Provide a backup for all essential components

    The University will provide a backup means for all essential components such as primary site located at a safe distance or cloud solution depending on the critical system to recover operations

  4. Provide for regular and effective Testing of the plan

The University will provide a full test of the IT DRP to ensure that the recovery steps and operations is feasible, and any problems / issues are identified in a timely manner and resolved appropriately.