ReleaseEngineering/How To/Process release email: Difference between revisions

no edit summary
(update for added info in gecko.git email)
No edit summary
Line 9: Line 9:
|-
|-
|Subject  || collapse report || [[#Performance Metrics]]
|Subject  || collapse report || [[#Performance Metrics]]
|-
|Subject  || Humpty Dumpty Error * || [[#Puppet failing too many times on a slave]]
|-
|Subject  || idle kittens report || [[#briar-patch idle kittens reporting]]
|-
|Subject ||[puppet-monitoring]* ||[[#Puppet Log Monitoring]]
|-
|-
|Subject || Suspected machine issue (* || Not an actionable email at this point. (from: nobody@cruncher - s/a {{bug|825625}}
|Subject || Suspected machine issue (* || Not an actionable email at this point. (from: nobody@cruncher - s/a {{bug|825625}}
Line 41: Line 35:
|-
|-
|To      || release+update.b2g.o@mozilla.com || low disk space on dogfood update server, see {{bug|877224}}
|To      || release+update.b2g.o@mozilla.com || low disk space on dogfood update server, see {{bug|877224}}
|-
|To      || release+chromecast@mozilla.com || Developer account for Chromecast app support {{bug|1037018}}
|}
|}


Line 46: Line 42:


__TOC__
__TOC__
=briar-patch idle kittens reporting=
== Why we get them ==
Email report outlining the status of any host that has been flagged as "idle"
== What is sending them ==
A cron job that is running the kittenreaper.py task with the following parameters
  python kittenreaper.py -w 1 -e
It pulls the list of hosts to check from http://builddata.pub.build.mozilla.org/reports/slaves_needing_reboot.txt
== What to do when one is received ==
not sure yet, unless your buildduty - then you should be watching it
== Future plans ==
This will be replaced by the briar-patch dashboard
== How to best filter these emails ==
Filtering can be done by matching the subject line which will not change
=Puppet Log Monitoring=
== Why we get them ==
There are messages in the puppet master logs that indicate something is wrong with a slave or master.  Since we have no other master monitoring tools, we are defaulting to sending email.
== What is sending them ==
scl-production-puppet and soon all puppet masters have an instance of 'watch-puppet.py' running under screen as root.
The code for this script is stored [https://github.com/jhford/monitor-puppet here]
== What to do when one is received ==
* if the title contains "[puppet-monitoring][master_name] <slavename> is waiting to be signed", this is for information and requires no immediate action
* if the title contains "[puppet-monitoring][master_name] <slavename> has invalid cert", the script will try once to clean the cert before sending the email once there is a waiting signing request.  If this is successful, you'll see a matching "<slavename> is waiting to be signed" email.  The key will be automatically signed by a cronjob
== How to silence or acknowledge this alert ==
It is not currently possible to silence this email.  This script will send email each time the corresponding line pattern is seen in /var/log/messages.  This means that most likely, each time a slave tries to puppet, an email will be sent.
== Future plans ==
In the short term, we'd like to have this script monitor the puppet logs for more error conditions.  It would also make sense to monitor all puppet masters
== How to best filter these emails ==
* subject includes [puppet-monitoring]
=Puppet failing too many times on a slave=
== Why we get them ==
We have no other monitoring for slaves failing to run puppet successfully.  This became a large issue with the rev4 talos machines due to {{bug|700672}}.  We are now doing an exponential back off on these slaves with a set number of iterations.  Once the maximum number of iterations is reached, the slave will send this email then reboot.  This helps us avoid puppet master load as well as allowing the machines try to fix themselves by rebooting.
== What is sending them ==
Each machine that has these emails enabled will send the email itself when it fails to puppet the last time, and right before it reboots.
The code that sends them is unversioned, but is deployed to the slaves from
scl-production-puppet:/N/production/darwin10-i386/test/usr/local/bin/run-puppet.sh
== What to do when one is received ==
* either ignore the email or find the root of the problem and fix it. 
== How to silence or acknowledge this alert ==
This email is a temporary workaround until we get a real puppet client monitoring tool.  This email we be sent each time the maximum number of retires is reached, which is every couple hours.
== Future plans ==
Would really like to replace these emails with real puppet monitoring.
== How to best filter these emails ==
These emails are best filtered by having "Humpty Dumpty Error" in their subject.  Becuase the hostname on the slave might not be correct every/all the time, filtering on domain names might not catch all cases.


=Performance Metrics=
=Performance Metrics=
Line 204: Line 136:


== How to silence or acknowledge this alert ==
== How to silence or acknowledge this alert ==


== Future plans ==
== Future plans ==


== How to best filter these emails ==
== How to best filter these emails ==
Filter on the sender and subject line.
Filter on the sender and subject line.


= Mail to release+chromecast@mozilla.com =
== Why we get them ==
The mobile team is adding Chromecast support (ability to fling videos/tabs from a device to a TV). They need a persistent account not linked to a single developer who might leave the company at some point.
== What is sending them ==
These emails come from the
== What to do when one is received ==
Traffic should be light. If the email is not simply Google self-promotion, please forward it to lead mobile devs, namely :blassey and :mfinkle.
== How to silence or acknowledge this alert ==
== Future plans ==
== How to best filter these emails ==
You can either filter on the "To:" field for "release+chromecast@mozilla.com" to catch just these emails, or filter on "From:" for "noreply@google.com" and move all mail from Google (we have multiple accounts mailing us intermittently) to a separate Google subfolder (coop).


<hr />
<hr />
canmove, Confirmed users
2,850

edits