<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0"><channel><title>SRE架构</title><link>https://sreai.net/zh/categories/sre%E6%9E%B6%E6%9E%84/</link><description>聚焦 SRE 与 AI，记录一个技术人的探索与沉淀</description><language>zh</language><lastBuildDate>Thu, 13 Aug 2026 23:52:33 +0800</lastBuildDate><item><title>深度学习在 SRE 智能运维中的实践探索</title><link>https://sreai.net/zh/posts/aiops-practice/</link><guid isPermaLink="true">https://sreai.net/zh/posts/aiops-practice/</guid><pubDate>Mon, 20 Jul 2026 10:00:00 +0800</pubDate><description>随着微服务架构与云原生规模的扩大，传统基于阈值的告警已无法满足快速止血需求。本文探讨了基于深度学习的智能异常检测系统如何帮助 SRE 团队从被动响应转向主动预测。</description><category>SRE架构</category><category>SRE</category><category>AIOps</category><category>机器学习</category></item><item><title>云原生混沌工程落地指南</title><link>https://sreai.net/zh/posts/chaos-engineering/</link><guid isPermaLink="true">https://sreai.net/zh/posts/chaos-engineering/</guid><pubDate>Thu, 18 Jun 2026 10:00:00 +0800</pubDate><description>通过 Chaos Mesh 模拟网络延迟、节点宕机与 Pod 异常，持续检验系统韧性与自动化自愈机制。</description><category>SRE架构</category><category>SRE</category><category>混沌工程</category><category>Kubernetes</category></item><item><title>Prometheus + eBPF 构建毫秒级可观测系统</title><link>https://sreai.net/zh/posts/ebpf-observability/</link><guid isPermaLink="true">https://sreai.net/zh/posts/ebpf-observability/</guid><pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate><description>结合 eBPF 零-overhead 内核探针与 Prometheus 联邦集群，构建覆盖网络、存储、应用三层的毫秒级全栈可观测体系。</description><category>SRE架构</category><category>eBPF</category><category>Prometheus</category><category>可观测性</category></item><item><title>Kubernetes 集群高可用与故障自动愈合实践</title><link>https://sreai.net/zh/posts/kubernetes-ha/</link><guid isPermaLink="true">https://sreai.net/zh/posts/kubernetes-ha/</guid><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><description>深入探讨控制面多活部署、etcd 集群调优、Pod 反亲和策略与自定义 Controller 实现故障自愈的最佳实践。</description><category>SRE架构</category><category>Kubernetes</category><category>SRE</category><category>高可用</category></item><item><title>GitOps 工作流在 Kubernetes 集群管理中的最佳实践</title><link>https://sreai.net/zh/posts/gitops-workflow/</link><guid isPermaLink="true">https://sreai.net/zh/posts/gitops-workflow/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>探索 GitOps 理念在 Kubernetes 集群管理中的落地实践，通过 ArgoCD 实现声明式部署与自动化回滚，提升发布效率与运维稳定性。</description><category>SRE架构</category><category>GitOps</category><category>Kubernetes</category><category>CI/CD</category></item><item><title>Terraform 基础设施即代码最佳实践</title><link>https://sreai.net/zh/posts/terraform-best-practices/</link><guid isPermaLink="true">https://sreai.net/zh/posts/terraform-best-practices/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>掌握 Terraform 基础设施即代码的核心实践，包括状态管理、模块化设计、远程后端配置和 CI/CD 集成</description><category>SRE架构</category><category>Terraform</category><category>IaC</category><category>DevOps</category></item><item><title>Envoy 服务网格流量管理实战</title><link>https://sreai.net/zh/posts/envoy-service-mesh/</link><guid isPermaLink="true">https://sreai.net/zh/posts/envoy-service-mesh/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>基于 Envoy 代理构建服务网格，实现高级流量管理、灰度发布和可观测性</description><category>SRE架构</category><category>Envoy</category><category>服务网格</category><category>可观测性</category></item><item><title>Vector：下一代可观测性数据管道</title><link>https://sreai.net/zh/posts/vector-observability/</link><guid isPermaLink="true">https://sreai.net/zh/posts/vector-observability/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>深入理解 Vector 数据管道的架构设计、高性能数据传输和统一可观测性策略</description><category>SRE架构</category><category>Vector</category><category>可观测性</category><category>日志</category></item></channel></rss>