‹ 返回事件历史

带有行动的事件

原文Incident with Actions
已恢复严重故障GitHub
2026年8月26日星期三 23:11 ~ 2026年8月27日星期四 02:01(2 小时 49 分钟)
受影响组件
Actions
Pages
更新记录
已恢复2026年8月26日星期三 18:01

2026年8月26日15:02至15:45 UTC期间,Actions作业无法启动。接下来的2小时直至17:40 UTC,由于系统追赶延迟负载,Actions运行启动延迟超过5分钟。此影响由处理Actions工作流触发器的服务所使用的数据库主库写入饱和所致。主库已进行故障转移,但系统未完全恢复。饱和现象源于日益增长的每日峰值负载,加上GitHub事件处理基础设施的上游问题,https://www.githubstatus.com/incidents/hcbtzksccj2f,导致本已高企的负载突发放大。后续用于恢复的下游限流设置高出约10%,未能有效保护系统。<br /><br />15:45 UTC时,限流与服务重启相结合恢复了服务的核心健康状态。这些限流在15:54至17:22之间逐步放宽,以恢复Actions运行的完整webhook处理。此提升过程刻意放缓,以确保在已知原始限流设置不当的情况下,不会再次压垮系统。webhook事件队列于17:40 UTC完全清空。<br /><br />3.7%的大型运行器作业以及部分规模集自托管作业仍停留在排队或“等待运行器”状态。我们部署了一项变更以强制撤销处于此状态的作业,并于18:40 UTC(事件缓解后约50分钟)将其转为失败状态。释放这些作业也同时释放了大型运行器作业的托管并发额度。<br /><br />使用并发组的客户受到更长时间的影响,原因是一个独立问题:在强制撤销缓解措施部署前,分配给部分作业子集的运行器断开连接,阻碍了运行器获取的进展,使作业停留在等待运行器状态。此问题于8月27日01:00 UTC解决。<br /><br />在15:02-15:45 UTC事件窗口期间触发的一些运行遇到了一个错误,导致它们在服务恢复后仍显示为排队状态。在后端,这些运行实际上已经失败,并将在创建24小时后自动转为取消状态。作为后续措施,我们正在修复此排队状态的根本原因,并提升批量取消受影响运行的能力。<br /><br />多项旨在提升Actions此部分整体可扩展性的变更已完成并正在部署至生产环境。这些变更的推出将在未来24小时内完成。进一步改善Actions工作流扩展性、弹性及更优雅降级的后续工作正在进行中。我们还将采取一项修复措施,以加速在类似未来情况下清除卡住的排队或等待作业。

原文On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. <br /><br />At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. <br /><br />3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. <br /><br />Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27. <br /><br />Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs. <br /><br />Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.

监控中2026年8月26日星期三 18:00

所有入站队列已恢复,Actions 运行正常。事件初期分配给较大运行器的 3.7% 任务仍卡在等待运行器分配状态,这些任务将在一小时内取消。其他运行器正在成功处理所有新任务。

原文All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.

监控中2026年8月26日星期三 17:54

影响Actions的服务降级问题已得到缓解。我们正在监控以确保稳定性。

原文The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

排查中2026年8月26日星期三 17:32

我们正在持续观察恢复情况,预计入站队列将在30分钟内恢复正常。工作将继续按每个客户的并发限制在系统中流转。

原文We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.

排查中2026年8月26日星期三 16:50

我们正在持续观察恢复情况,延迟队列正在逐步消化。在受限工作全部完成之前,部分客户仍会看到延迟增加——预计将在接下来的一小时内完成。

原文We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.

排查中2026年8月26日星期三 16:49

Pages 运行正常。

原文Pages is operating normally.

排查中2026年8月26日星期三 16:14

我们相信已经定位并解决了问题,正在逐步恢复流量,以确保问题不再发生。随着流量的逐步恢复,部分客户可能仍会看到延迟现象。

原文We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

排查中2026年8月26日星期三 15:48

主故障切换短暂提升了性能,但并未完全缓解问题,我们已限制入站流量,并正在调查上游Vitess相关问题。

原文primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

排查中2026年8月26日星期三 15:23

我们已发现数据库主实例存在问题,正在立即故障转移至副本。

原文We've identified an issue with a database primary and are failing over to a replica immediately

排查中2026年8月26日星期三 15:12

Pages 正在经历性能降级。我们正在继续调查。

原文Pages is experiencing degraded performance. We are continuing to investigate.

排查中2026年8月26日星期三 15:11

我们正在调查有关Actions可用性下降的报告。

原文We are investigating reports of degraded availability for Actions

事件内容来自 GitHub 官方状态页:查看官方原文。中文由 AI 翻译,仅供快速理解,以官方英文原文为准。

关于这些数据

GitHub的故障事件数据从哪来?

全部来自GitHub 官方状态页的公开接口,由本站每 5 分钟同步一次,保留最近 90 天。事件标题、时间、影响级别与更新记录均为官方原文,本站不做改写;点详情页底部的链接可回到厂商原始记录核对。

顶部的组件状态条怎么看?

每一格是一天,绿色表示GitHub 官方状态页当天未记录异常,黄色表示部分降级,红色表示当天有较大故障,灰色表示该组件当时还没有数据。右侧的可用率是官方口径的 90 天统计,与状态总览页显示的是同一份数据。

为什么有的服务查不到历史?

本页只收录对外提供官方状态页的服务。没有公开状态页的厂商(多数国产大模型属于此类)无法取得可信数据,本站不做自建拨测去猜,因此也不会出现在状态总览里。

持续时长怎么算?

按官方标注的开始时间到恢复时间计算;尚未恢复的事件按「至今」计算并标为进行中。跨天的事件在状态条上会覆盖它经过的每一天。

完整的服务清单与实时状态见状态总览