http://subversion.tigris.org/issues/show_bug.cgi?id=1964
The issue begin at 8:00PM PST time. The issue was briefly resolved at 9:30PM PST time but the site experienced another outage starting at 9:50PM PST due to Edgecasts operators insistance that there was a hardware issue with the system which is what caused it to become unresponse (even to DRAC).
At approximately 12:30AM we remounted the site on our development systems until Edgecast was able to restore the downed NFS server. at 1:20AM.
Action Items:
- Take the snapshot of partner source code and do an svn import instead of an svn add on the development server. We need this code in source control to ensure that we can better support our development process. This avoids the issue: http://subversion.tigris.org/issues/show_bug.cgi?id=1964
- Create a status blog that we can give partners so that they can log in and track our progress on site issues.
- Collect all partner emails so that we can immediately inform them of site outages as they occur.
- Determine how we can better isolate our traffic from partner traffic so that it does not affect the partner when we go down. This will likely be a combination of using a CDN as well as a maintenance site/httpd server that will serve up a “currently down for maintenance” page and also will quickly return/404 flash calls to the flash swf that is embedded for partner tracking/reporting.
- Move away from the NFS mount for code hosting or at least use DRDB replication to ensure high availability on the NFS mount. The NFS mount is the biggest SPOF in our system and we need to address it VERY soon to ensure our ability to tolerate failures and avoid site outages.
No comments:
Post a Comment