<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0"><channel><title>SRE</title><link>https://sreai.net/en/tags/sre/</link><description>Practical notes on SRE, cloud-native reliability, and AI infrastructure.</description><language>en</language><lastBuildDate>Thu, 13 Aug 2026 23:52:33 +0800</lastBuildDate><item><title>Deep Learning in SRE Intelligent Operations Practice</title><link>https://sreai.net/en/posts/aiops-practice/</link><guid isPermaLink="true">https://sreai.net/en/posts/aiops-practice/</guid><pubDate>Mon, 20 Jul 2026 00:00:00 +0000</pubDate><description>As microservice architectures and cloud-native scale grow, traditional threshold-based alerting can no longer meet the need for rapid incident response. This article explores how deep learning-based anomaly detection systems help SRE teams shift from reactive responses to proactive prediction.</description><category>SRE Architecture</category><category>SRE</category><category>AIOps</category><category>机器学习</category></item><item><title>Cloud-Native Chaos Engineering Practice Guide</title><link>https://sreai.net/en/posts/chaos-engineering/</link><guid isPermaLink="true">https://sreai.net/en/posts/chaos-engineering/</guid><pubDate>Thu, 18 Jun 2026 00:00:00 +0000</pubDate><description>Use Chaos Mesh to simulate network latency, node failures, and Pod anomalies to continuously verify system resilience and automated self-healing mechanisms.</description><category>SRE Architecture</category><category>SRE</category><category>混沌工程</category><category>Kubernetes</category></item><item><title>Kubernetes High Availability and Fault Auto-Healing Practice</title><link>https://sreai.net/en/posts/kubernetes-ha/</link><guid isPermaLink="true">https://sreai.net/en/posts/kubernetes-ha/</guid><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><description>Deep dive into control plane multi-active deployment, etcd cluster tuning, Pod anti-affinity strategies, and custom Controller-based fault self-healing best practices.</description><category>SRE Architecture</category><category>Kubernetes</category><category>SRE</category><category>High-Availability</category></item><item><title>Building an Effective On-Call System: From Postmortems to Proactive Defense</title><link>https://sreai.net/en/posts/incident-management/</link><guid isPermaLink="true">https://sreai.net/en/posts/incident-management/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>From alert design, shift scheduling to incident response workflows, systematically build an effective On-Call system that shifts from reactive firefighting to proactive defense.</description><category>Technology Management</category><category>On-Call</category><category>Management</category><category>SRE</category></item></channel></rss>