Platform Operations Support Role
Work in progress - In these early stages of the Platform Operations team, we are still defining the support role and responsibilities. This page will be updated as we continue to refine our support processes.
Allocating Engineers for Support
We have a daily support rota for the Platform Operations team, with both Primary and Secondary roles.
- Primary Engineer: The primary point of contact for support on a given day - they should be the first to respond to support queries/alerts etc
- Secondary Engineer: Can be used as a point of escalation/support for the primary engineer, and may also cover for the Primary Engineer if they are temporarily unavailable
These roles are responsible for:
- Acknowledging and triaging support requests
- Monitoring the Slack alert channels
- Monitoring PagerDuty for any health issues or incidents.
A daily Github Actions build will detect the allocated Primary/Secondary Engineers for the day, and will leverage the #internal-platform-operations channel to notify us of who they are.
If you are the Primary or Secondary Engineer for the day, please respond with a
:white_tick: emoji.
We use Slack emojis to indicate our assigned Support Role status. Primary is
:platops-platypus:, Secondary is
:octopus:.
If you are the allocated engineer, but you are unable to fulfill the role for the day, please send a message to the #internal-platform-operations channel (tagging @platform-operations) and request that someone else takes on the role in your place.
If the allocated Primary Engineer for the day has to change, please ensure that an override is added to the daily support rota, so that the new Primary Engineer is reflected, as this will ensure that they receive any relevant alerts throughout the course of the day.
Our general support pattern is from 09:00 to 16:00 (though it should be noted that our PagerDuty rota runs from 08:00 to 16:00) - we do not currently have an on-call capability.
Daily Concierge duties
Below are some common support tasks that you would be expected to carry out throughout the course of the day:
- NDST Standup: The Primary and Secondary Engineers should join the NDST Standup at 10:00am. This allows us to communicate with the relevant stakeholders, provide status updates and respond to any queries that may arise.
- NOTE: You do not need to stay for the full duration of the call - once the DRB review begins, you can leave.
- Monitor Support Channels: You should monitor the public support channels throughout the day, responding to any queries that may arise.
- Montior PagerDuty: You should monitor PagerDuty throughout the day, responding to any alerts.
- Our core support hours run from 9am to 4pm Monday to Friday. Support outside of this hours is on a “best endeavours” basis.
- It is also recommended that, at the start of the day, you look at all currently open incidents and filter by the OCTO Platform Operations team - this will allow you to quickly see if there are any alerts from the previous day (or that may have triggered overnight) that have not yet been escalated.
Dependabot PR Reviews: Check for any outstanding Dependabot PRs on the following repositories:
DSO Infra Azure AD User Management: On Tuesday mornings, an automated-user-management will run - this detects any users who should be marked as “inactive”, or any users who should be “deleted”, and will generate PRs if appropriate. If today is a Tuesday, go to the dso-infra-azure-ad, ensure that the workflow has run successfully, and if there are any PRs that have been generated, review them.
As it pertains to DSO services, supporting guidance for being on support can be found below:
- Common Concierge Tasks
- Concierge Tasks
- DSO Self Service Guides
- Handling PagerDuty Alarms
- Monitoring and Alerting
Note: It is not expected that support team members will always be able to resolve all issues, but they should respond, and escalate issues to the appropriate team members where necessary.
Support tickets
Some support requests or platform alerts may require a considerable investment of time to close. Some example scenarios include:
A complex change request that requires significant review time and/or testing, or coordination with other teams.
An issue or request that is likely to remain open for more than a day, to allow handover to be documented and for the next support team member to pick up the ticket.
A routine, regular task that has not yet been either automated or added to the runbooks, and requires a support team member to perform the task manually. Documenting the process should be covered as part of the created ticket.
If you want to add a support ticket in the Platform Operations Jira board, create it under parent EPIC PlatOps Support Rota & Ad-Hoc Requests and assign/place in sprint appropriately.
Alerting Slack Channels
Whilst on support, you should keep an eye on the following Slack channels, which are used for monitoring and alerting:
Non-Prod Alarm Slack Channels
Contains alerts from AWS modernisation platform development, test and pre-production accounts.
#az_noms_dev_test_environments_alerts- For DSO DevTest Azure alerts#dso_alerts_pipeline- For scheduled tasks or pipeline failures within the DSO estate#dba_alerts_devtest- For non-production DBA alerts within the DSO Estate
The following are non-production alerting channels for projects that DSO supports:
#hmpps_domain_services_alarms_non_prod(shared AD/Remote Desktop resources)#hmpps_oem_alarms_non_prod#nomis_alarms_non_prod#nomis_combined_reporting_alarms_non_prod#nomis_data_hub_alarms_non_prod#oasys_alarms_non_prod#oasys_national_reporting_alarms_non_prod
Production Alarm Slack Channels
Contains alerts from AWS modernisation platform production accounts and NOMS Production Azure Account.
#az_noms_production_1_alerts- For DSO production Azure alerts#az_noms_notifications- Receives notifications for NOMS production (Azure)#corporate_staff_rostering_alarms- Alarms for CSR#dba_alerts_prod- For production DBA alerts within the DSO Estate
The following are production alerting channels for projects that DSO supports:
#hmpps_domain_services_alarms_prod(shared AD/Remote Desktop resources)#hmpps_oem_alarms_prod#nomis_alarms_prod#nomis_combined_reporting_alarms_prod#nomis_data_hub_alarms_prod#oasys_alarms_prod#oasys_national_reporting_alarms_prod#planetfm_alarms#prison_retail_alarms
Daily Security Hub Notification
Within the #dso_alerts_pipeline channel, we receive daily notifications of any new Critical/High Security Hub alerts for projects/environments in AWS that we manage. These are generated via the following Github Workflow.
You should keep an eye on this and review as appropriate
Email Notifications
Occassionally, we will receive notifications via email to our Digital Studio Operations and Team Migrations inboxes - examples of this include:
- AWS Health Notifications (e.g. service outages, maintenance notifications)
- Azure Notifications (e.g. security recommendations, outages, weekly digests)
Please ensure that you monitor emails throughout the day, and respond as appropriate. If you are unsure of how to proceed following an email, if there is a major service outage, or if you suspect you do not have access to these email inboxes, please contact
the team by sending a message to the #internal-platform-operations channel (tagging @platform-operations)