‹ 返回事件历史

带有行动的事件

原文Incident with Actions
已恢复轻微故障GitHub
2026年7月25日星期六 16:59 ~ 2026年7月25日星期六 17:25(25 分钟)
受影响组件
Actions
更新记录
已恢复2026年7月25日星期六 09:25

2026年7月25日,GitHub Actions经历了两个相关的性能下降时段,导致部分工作流运行延迟超过5分钟或最终因基础设施故障而失败。<br /><br />第一时段(08:45 – 09:13 UTC):在对Actions的关键路径Redis集群进行计划维护期间,一个参与区域处于降级状态。另外,一次独立的容量操作暂时将另一个区域从集群中移除,并将其流量重定向至降级区域。这导致了作业分配状态在跨区域间出现不一致,使得工作流运行延迟、重试耗尽或直接失败。在高峰期,约7%的运行延迟超过5分钟,事件期间有25%的运行因基础设施错误而失败。我们于09:13 UTC通过将流量恢复正常分布来缓解了此事件。<br /><br />第二时段(12:08 – 12:48 UTC):作为缓解第一事件的一部分,流量被恢复到仍在进行容量扩展的区域实例。扩展区域中的多个Redis节点发生故障,增加了健康节点的流量,导致许多节点达到连接限制。在高峰期,30%的运行延迟超过5分钟,事件期间有60%的运行因基础设施错误而失败。我们于12:48 UTC通过将工作流流量从扩展区域转移来缓解了此事件。<br /><br />我们正在增加更严格的区域健康和容量检查,在维护前执行,并要求在恢复流量前有稳定的观察期。我们还在改进自动连接弹性,并与我们的平台依赖方合作,自动检测和修复不健康的集群成员及分片不平衡问题。更广泛地说,我们已经在进行中的工作旨在提高Actions基础设施这一部分的弹性和规模。

原文On July 25, 2026, GitHub Actions experienced two related periods of degradation that caused some workflow runs to be delayed by more than 5 minutes or end with infrastructure failures. <br /><br />First period (08:45 – 09:13 UTC): During planned maintenance on a critical-path Redis cluster for Actions, one participating region was left in a degraded state. Separately, an independent capacity operation temporarily removed another region from the cluster and redirected its traffic to the degraded region. This created cross-region inconsistencies in job-assignment state, causing workflow runs to be delayed, exhaust retries, or fail outright. At peak, about 7% of runs were delayed by more than 5 minutes, and 25% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 09:13 UTC by returning traffic to its normal distribution. <br /><br />Second period (12:08 – 12:48 UTC): As part of mitigating the first incident, traffic was returned to the regional instance that was still undergoing its capacity increase. Multiple Redis nodes in the scaling region experienced failures, increasing traffic to healthy nodes and causing connection limits to be reached on many nodes. At peak, 30% of runs were delayed by more than 5 minutes, and 60% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 12:48 UTC by redirecting workflow traffic away from the scaling region. <br /><br />We are adding stronger regional health and capacity checks before maintenance and requiring a stable observation period before restoring traffic. We are also improving automated connection resiliency, and partnering with our platform dependency to automatically detect and remediate unhealthy cluster members and shard imbalance. More generally, we already had work underway to improve the resiliency and scale of this piece of Actions infrastructure.

监控中2026年7月25日星期六 09:20

我们已识别出导致GitHub Actions运行启动延迟的问题。部分用户在触发工作流运行时可能遇到了比预期更长的等待时间。我们已采取缓解措施并已恢复。团队正在持续监控并调查根本原因。

原文We identified an issue causing delays in GitHub Actions run starts. Some users may have experienced longer than expected wait times when triggering workflow runs. We have applied mitigations and have recovered. Our team continues to monitor and investigate the root cause.

监控中2026年7月25日星期六 09:13

影响Actions的服务降级问题已得到缓解。我们正在监控以确保稳定性。

原文The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

排查中2026年7月25日星期六 08:59

我们正在调查有关Actions性能下降的报告。

原文We are investigating reports of degraded performance for Actions

事件内容来自 GitHub 官方状态页:查看官方原文。中文由 AI 翻译,仅供快速理解,以官方英文原文为准。

关于这些数据

事件从哪来?

全部来自各厂商官方状态页的公开接口,由本站每 5 分钟同步一次,保留最近 90 天。事件标题、时间、影响级别与更新记录均为官方原文,本站不做改写;点详情页底部的链接可回到厂商原始记录核对。

为什么有的服务查不到历史?

本页只收录对外提供官方状态页的服务。没有公开状态页的厂商(多数国产大模型属于此类)无法取得可信数据,本站不做自建拨测去猜,因此也不会出现在状态总览里。

持续时长怎么算?

按官方标注的开始时间到恢复时间计算;尚未恢复的事件按「至今」计算并标为进行中。跨天的事件在状态条上会覆盖它经过的每一天。