GitHub的一些服务中断
7月24日16:04 UTC,我们的三个物理数据中心可用区(AZ)之一中的网络路径发生连接中断。这导致剩余活动路径饱和,从而引发数据包丢失。我们的数据中心在每个计算笼中使用叶脊交换机架构,并在每个AZ内通过聚合层互连各笼的脊交换机。连接中断影响了该特定AZ内一个笼的脊交换机与聚合层之间的链路。<br /><br />依赖此笼中计算资源的工作负载因数据包丢失而性能下降,并出现间歇性错误:<br /><br />- 操作在影响窗口期间有10%的作业失败,5%的作业成功但启动延迟。<br />- 27%的GitHub问题交互遇到慢请求或超时。<br />- 4%的GitHub Copilot请求出现错误,尽管大多数会自动重试。<br />- 受影响窗口期间,4%的git推送操作受到影响。<br />- 认证请求在受影响窗口期间延迟增加,但错误率虽有所上升,在所有情况下均低于1%。<br /><br />我们通过将受影响的连接重新路由到为未来容量升级分配的可用光纤路径,成功缓解了此次中断。17:07恢复了足够的网络容量以消除数据包丢失,大多数服务在17:16完全恢复。所有路径于17:36恢复,服务健康。<br /><br />此次事件影响了可用网络互连容量的25%。较旧的笼采用100Gbps网络接口标准。为消除再次发生的风险,计划中的400Gbps接口升级正在尽可能加速推进,确保交换机架构各层带宽增加,以增强对路径或设备丢失的恢复能力。
原文On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ. <br /><br />Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors: <br /><br />- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts. <br />- 27% of GitHub issues interactions saw slow requests or timeouts. <br />- 4% of GitHub Copilot requests experienced errors, though most automatically retry. <br />- 4% of git push operations saw impacts during the affected window. <br />- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases. <br /><br />We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36. <br /><br />This incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss.
我们正在看到所有服务的恢复。
原文We are seeing recovery across all services
影响API请求、Actions、Copilot、Issues、Pages和Pull Requests的性能降级问题已得到缓解。我们正在持续监控以确保稳定性。
原文The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability.
Actions 正在经历性能降级。我们正在继续调查。
原文Actions is experiencing degraded performance. We are continuing to investigate.
我们已采取缓解措施,并正在监控恢复情况。
原文We have applied a mitigation and are monitoring for recovery
Actions 正经历可用性降级。我们正在继续调查。
原文Actions is experiencing degraded availability. We are continuing to investigate.
Pages 正在经历性能降级。我们正在继续调查。
原文Pages is experiencing degraded performance. We are continuing to investigate.
Copilot 正在经历性能下降。我们正在继续调查。
原文Copilot is experiencing degraded performance. We are continuing to investigate.
我们正在调查部分GitHub服务的超时问题。
原文We are investigating timeouts to some GitHub services
Pull Requests 正经历性能下降。我们正在继续调查。
原文Pull Requests is experiencing degraded performance. We are continuing to investigate.
Actions 正在经历性能降级。我们正在继续调查。
原文Actions is experiencing degraded performance. We are continuing to investigate.
我们正在调查有关API请求和问题性能下降的报告。
原文We are investigating reports of degraded performance for API Requests and Issues
关于这些数据
事件从哪来?
全部来自各厂商官方状态页的公开接口,由本站每 5 分钟同步一次,保留最近 90 天。事件标题、时间、影响级别与更新记录均为官方原文,本站不做改写;点详情页底部的链接可回到厂商原始记录核对。
为什么有的服务查不到历史?
本页只收录对外提供官方状态页的服务。没有公开状态页的厂商(多数国产大模型属于此类)无法取得可信数据,本站不做自建拨测去猜,因此也不会出现在状态总览里。
持续时长怎么算?
按官方标注的开始时间到恢复时间计算;尚未恢复的事件按「至今」计算并标为进行中。跨天的事件在状态条上会覆盖它经过的每一天。