CIDuty: Difference between revisions
ChrisCooper (talk | contribs) |
ChrisCooper (talk | contribs) |
||
| Line 106: | Line 106: | ||
= Standard Bugs = | = Standard Bugs = | ||
* For IT bugs that are marked "infra only", yet still need to be readable by RelEng, it is not enough to add release@ alias - people get updates but not able to comment or read prior comments. Instead, cc the following: | * For IT bugs that are marked "infra only", yet still need to be readable by RelEng, it is not enough to add release@ alias - people get updates but not able to comment or read prior comments. Instead, cc the following: | ||
** : | ** :bhearsum, :Callek, :catlee, :coop, :hwine, :jlund, :kmoir, :mrrrgn, :nthomas, :rail | ||
** | ** :Tomcat, :RyanVM, :KWierso | ||
= Meeting Notes = | = Meeting Notes = | ||
* [https://releng.etherpad.mozilla.org/buildduty buildduty log] | * [https://releng.etherpad.mozilla.org/buildduty buildduty log] | ||
* [[ReleaseEngineering/Buildduty/Meetings|Meetings Notes]] | * [[ReleaseEngineering/Buildduty/Meetings|Meetings Notes]] | ||
Revision as of 17:16, 1 April 2015
What is buildduty?
Every month, there is one person from the Release Engineering (releng) team dedicated to helping out developers with releng-related issues. This person will be available during his or her regular work hours for the whole month. This is similar to the sheriff role that rotates through the sheriffing team . To avoid confusion, the releng sheriff position is known as "buildduty."
Who is on buildduty? (schedule)
The person on buildduty should have 'buildduty' appended to their IRC nick, and should be available in the #developers, #releng, and #buildduty IRC channels.
Mozilla Releng Buildduty Schedule (Google Calendar|iCal|XML)
Buildduty not around?
It happens, especially outside of standard North American working hours (0600-1800 PT). Please open a bug under these circumstances.
Buildduty priorities
How should I make myself available for duty?
- Add 'buildduty' to your IRC nick
- Be available in the following IRC channels (at least): #developers, #releng, and #buildduty (as well as #mozbuild of course)
- also useful to be in #mobile and #ateam
- if you are in the middle of an outage, or need IT help, it is useful to be in #moc, #infra, and #sysadmins.
What should I take care of?
Outages
Things fail. It's sad. Getting systems and services stood back up again is buildduty's top priority. Note: this doesn't mean you need to do all the work yourself. For big outages, rope in whatever help you need: domain experts from releng, managers, netops, relops...whoever.
Daily
Triage
The Buildduty report (generated hourly) should be your starting point for triage.
At the top, it lists unassigned bugs for loan requests. You should try to keep this queue empty to make sure developers are unblocked. The wiki has instructions for how to loan a slave.
After loans are taken care of, make sure that bugs in the "No dependencies" section get dependencies filed, e.g. diagnosis bug, decomm bug, etc. Do the same for bugs in the "All dependencies resolved" section to make sure the next action is taken (re-image, decomm, return to production, etc). Use the "View list in bugzilla" links to navigate the bugs more easily.
Note: systemic issues (e.g. test failures that require further investigation) should *not* stay in the buildduty bugzilla component. It may be OK for you to take the bug and work on it depending on how much time you have, but generally these types of bugs should be moved to a more-appropriate component (e.g. General Automation) once buildduty has triaged them.
Aside from the buildduty report, there may also be unacknowledged nagios alerts in the #buildduty IRC channel. Deal with them, filing bugs as needed. |}
Semi-Daily
- Reconfigs
- Review "long running" and "lazy" AWS instances
- When: 2-3 times a week (eg: Mondays or after weekends/holidays, Wednesdays, and Fridays)
- How: use aws sanity check email (sent daily):
- Email filter => to: release+aws-sanity-check@mozilla.com, subject: [cron] aws sanity check
- for each host under heading "Long running instances", follow steps in dealing with long running instances
- for each host under heading "Lazy long running instances", figure out why they're still up and not taking jobs
- twistd.log, uptime, reboot history on the slave health page et al
Weekly
- Review AWS instances that have 'Unknown State/Type or have stopped for a while'
- When:
- Once a week. Preferably, this would be evenly spaced out between tackling this so let's say Fridays if possible.
- How:
- use aws sanity check email (sent daily):
- Email filter => to: release+aws-sanity-check@mozilla.com, subject: [cron] aws sanity check
- for each host under heading "Unknown State", "Unknown Type"
- follow steps in dealing with unknown state or type instances
- for each host under heading "Stopped For A While"
- follow steps in dealing with stopped for a while instances
- use aws sanity check email (sent daily):
- When:
Infrastructure performance
Pending Jobs
We will sometimes be starved for capacity on one or more platforms. Because there are multiple potential causes, and hence multiple possible paths to resolution, the steps for dealing with high pending counts are on their own page.
Wait times
This can be related to pending builds above.
Releng has made a commitment to developers that 95% or more of their jobs will start within 15 minutes of submission.
Build and Try (Build) slave pools have greater capacity (and can expand into AWS as required for linux/mobile/b2g) and are usually over 95% unless there is an outage.
Many Test jobs are triggered per build/try job, and the current slave pool is finite, so it is rare for us to meet our turnaround commitment for test jobs.
Fixing errant test slaves is hence more important fixing build slaves. See Slave Management below.
Wait times are available either from the buildAPI wait times report or the daily emails that go to dev.tree-management (un-filter them in Zimbra). Respond to any unusually long wait times in email, preferably with a reason.
Wait times emails are run via crontab entries setup on relengwebadm.private.scl3.mozilla.com under the buildapi user.
Slave management
Bad slaves can burn builds and hung slaves can cause bad wait times. These slaves need to be rebooted, or handed to IT for recovery. Recovered slaves need to be tracked on their way back to operational status.
The Nagios wiki has more information about finding problem slaves using nagios.
See the Slave Management wiki for more information about fixing those slaves.
dev.tree-management
Monitor dev.tree-management newsgroup (by email or by nntp).
Watch for long running builds that are holding on to slave, i.e. >1 day.
See the buildAPI list of running builds.
Others
You should keep on top of:
- Developer requests in IRC
- Direct people to http://mzl.la/tryhelp for self-serve documentation when appropriate.
- Other less-frequent duties
Useful Links
- Slave Health
- Build Dashboard Main Page
- You can get JSON dumps for people to analyze by adding
&format=json
- You can get JSON dumps for people to analyze by adding
- Public "How To" documents
- Private "How To" documents
Standard Bugs
- For IT bugs that are marked "infra only", yet still need to be readable by RelEng, it is not enough to add release@ alias - people get updates but not able to comment or read prior comments. Instead, cc the following:
- :bhearsum, :Callek, :catlee, :coop, :hwine, :jlund, :kmoir, :mrrrgn, :nthomas, :rail
- :Tomcat, :RyanVM, :KWierso