<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0"><channel><title>博文</title><link>https://sreai.net/zh/posts/</link><description>聚焦 SRE 与 AI，记录一个技术人的探索与沉淀</description><language>zh</language><lastBuildDate>Thu, 13 Aug 2026 23:52:33 +0800</lastBuildDate><item><title>深度学习在 SRE 智能运维中的实践探索</title><link>https://sreai.net/zh/posts/aiops-practice/</link><guid isPermaLink="true">https://sreai.net/zh/posts/aiops-practice/</guid><pubDate>Mon, 20 Jul 2026 10:00:00 +0800</pubDate><description>随着微服务架构与云原生规模的扩大，传统基于阈值的告警已无法满足快速止血需求。本文探讨了基于深度学习的智能异常检测系统如何帮助 SRE 团队从被动响应转向主动预测。</description><category>SRE架构</category><category>SRE</category><category>AIOps</category><category>机器学习</category></item><item><title>Go 语言微服务架构实践与性能调优</title><link>https://sreai.net/zh/posts/go-microservices/</link><guid isPermaLink="true">https://sreai.net/zh/posts/go-microservices/</guid><pubDate>Fri, 15 May 2026 10:00:00 +0800</pubDate><description>本文将深入探讨使用 Go 语言构建高并发微服务体系的具体实践，包括 API 契约设计、服务发现以及负载均衡策略。</description><category>编程开发</category><category>Go</category><category>微服务</category><category>API</category></item><item><title>云原生混沌工程落地指南</title><link>https://sreai.net/zh/posts/chaos-engineering/</link><guid isPermaLink="true">https://sreai.net/zh/posts/chaos-engineering/</guid><pubDate>Thu, 18 Jun 2026 10:00:00 +0800</pubDate><description>通过 Chaos Mesh 模拟网络延迟、节点宕机与 Pod 异常，持续检验系统韧性与自动化自愈机制。</description><category>SRE架构</category><category>SRE</category><category>混沌工程</category><category>Kubernetes</category></item><item><title>Prometheus + eBPF 构建毫秒级可观测系统</title><link>https://sreai.net/zh/posts/ebpf-observability/</link><guid isPermaLink="true">https://sreai.net/zh/posts/ebpf-observability/</guid><pubDate>Fri, 10 Jul 2026 00:00:00 +0000</pubDate><description>结合 eBPF 零-overhead 内核探针与 Prometheus 联邦集群，构建覆盖网络、存储、应用三层的毫秒级全栈可观测体系。</description><category>SRE架构</category><category>eBPF</category><category>Prometheus</category><category>可观测性</category></item><item><title>Kubernetes 集群高可用与故障自动愈合实践</title><link>https://sreai.net/zh/posts/kubernetes-ha/</link><guid isPermaLink="true">https://sreai.net/zh/posts/kubernetes-ha/</guid><pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate><description>深入探讨控制面多活部署、etcd 集群调优、Pod 反亲和策略与自定义 Controller 实现故障自愈的最佳实践。</description><category>SRE架构</category><category>Kubernetes</category><category>SRE</category><category>高可用</category></item><item><title>基于 RAG 构建企业级智能运维知识库</title><link>https://sreai.net/zh/posts/rag-knowledge-base/</link><guid isPermaLink="true">https://sreai.net/zh/posts/rag-knowledge-base/</guid><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><description>利用 LangChain + Chroma 向量数据库，将历史 Runbook、事故报告和运维文档转化为可检索的智能知识库，实现 On-Call 场景下的快速根因定位。</description><category>人工智能</category><category>RAG</category><category>AIOps</category><category>LLM</category></item><item><title>LLM Infra 搭建与 GPU 集群性能调优</title><link>https://sreai.net/zh/posts/llm-gpu-infra/</link><guid isPermaLink="true">https://sreai.net/zh/posts/llm-gpu-infra/</guid><pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate><description>从零搭建大语言模型推理基础设施，涵盖 vLLM 分布式部署、GPU 显存优化策略与 NVIDIA MIG 切分实践。</description><category>人工智能</category><category>LLM</category><category>GPU</category><category>AI</category></item><item><title>2026 技术人的架构演进思考</title><link>https://sreai.net/zh/posts/architecture-evolution-2026/</link><guid isPermaLink="true">https://sreai.net/zh/posts/architecture-evolution-2026/</guid><pubDate>Thu, 23 Jul 2026 00:00:00 +0000</pubDate><description>站在 2026 年中回望，从云原生到 AI-Native，从微服务到 Agent 协作，技术架构正在经历前所未有的范式转移。本文记录一个 SRE 工程师的观察与思考。</description><category>技术管理</category><category>架构</category><category>技术管理</category></item><item><title>GitOps 工作流在 Kubernetes 集群管理中的最佳实践</title><link>https://sreai.net/zh/posts/gitops-workflow/</link><guid isPermaLink="true">https://sreai.net/zh/posts/gitops-workflow/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>探索 GitOps 理念在 Kubernetes 集群管理中的落地实践，通过 ArgoCD 实现声明式部署与自动化回滚，提升发布效率与运维稳定性。</description><category>SRE架构</category><category>GitOps</category><category>Kubernetes</category><category>CI/CD</category></item><item><title>Python 异步编程模式：从协程到事件循环深度剖析</title><link>https://sreai.net/zh/posts/python-async-patterns/</link><guid isPermaLink="true">https://sreai.net/zh/posts/python-async-patterns/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>深入解析 Python asyncio 库的核心机制，从协程定义、事件循环调度到异步上下文管理，帮助开发者编写高性能并发代码。</description><category>编程开发</category><category>Python</category><category>异步</category><category>API</category></item><item><title>MLOps 流水线实战：模型训练到部署的自动化之路</title><link>https://sreai.net/zh/posts/mlops-pipeline/</link><guid isPermaLink="true">https://sreai.net/zh/posts/mlops-pipeline/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>从数据准备、模型训练到生产部署，构建端到端的 MLOps 自动化流水线，实现机器学习模型持续交付与监控。</description><category>人工智能</category><category>MLOps</category><category>AI</category><category>机器学习</category></item><item><title>分布式数据库分库分表策略与一致性保障</title><link>https://sreai.net/zh/posts/database-sharding/</link><guid isPermaLink="true">https://sreai.net/zh/posts/database-sharding/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>深入探讨分布式数据库场景下的分库分表设计策略，涵盖哈希分片、范围分片及分布式事务一致性保障方案。</description><category>编程开发</category><category>数据库</category><category>分布式</category><category>Go</category></item><item><title>构建高效 On-Call 体系：从事后复盘到主动防御</title><link>https://sreai.net/zh/posts/incident-management/</link><guid isPermaLink="true">https://sreai.net/zh/posts/incident-management/</guid><pubDate>Sat, 25 Jul 2026 00:00:00 +0000</pubDate><description>从告警策略设计、值班排班管理到故障响应流程，系统化构建高效的 On-Call 体系，实现从被动救火到主动防御的转变。</description><category>技术管理</category><category>On-Call</category><category>技术管理</category><category>SRE</category></item><item><title>Terraform 基础设施即代码最佳实践</title><link>https://sreai.net/zh/posts/terraform-best-practices/</link><guid isPermaLink="true">https://sreai.net/zh/posts/terraform-best-practices/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>掌握 Terraform 基础设施即代码的核心实践，包括状态管理、模块化设计、远程后端配置和 CI/CD 集成</description><category>SRE架构</category><category>Terraform</category><category>IaC</category><category>DevOps</category></item><item><title>gRPC-Gateway：构建高性能 API 网关</title><link>https://sreai.net/zh/posts/grpc-gateway/</link><guid isPermaLink="true">https://sreai.net/zh/posts/grpc-gateway/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>深入探讨 gRPC-Gateway 的架构设计、protobuf 注解配置和 REST API 转换实践</description><category>编程开发</category><category>gRPC</category><category>Go</category><category>微服务</category></item><item><title>Envoy 服务网格流量管理实战</title><link>https://sreai.net/zh/posts/envoy-service-mesh/</link><guid isPermaLink="true">https://sreai.net/zh/posts/envoy-service-mesh/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>基于 Envoy 代理构建服务网格，实现高级流量管理、灰度发布和可观测性</description><category>SRE架构</category><category>Envoy</category><category>服务网格</category><category>可观测性</category></item><item><title>用 Rust 构建高效 CLI 运维工具</title><link>https://sreai.net/zh/posts/rust-cli-tools/</link><guid isPermaLink="true">https://sreai.net/zh/posts/rust-cli-tools/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>利用 Rust 语言的安全性和高性能特性，构建现代化的命令行运维工具集</description><category>编程开发</category><category>Rust</category><category>CLI</category><category>运维</category></item><item><title>Vector：下一代可观测性数据管道</title><link>https://sreai.net/zh/posts/vector-observability/</link><guid isPermaLink="true">https://sreai.net/zh/posts/vector-observability/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>深入理解 Vector 数据管道的架构设计、高性能数据传输和统一可观测性策略</description><category>SRE架构</category><category>Vector</category><category>可观测性</category><category>日志</category></item><item><title>Kubernetes FinOps：云成本优化实践</title><link>https://sreai.net/zh/posts/finops-kubernetes/</link><guid isPermaLink="true">https://sreai.net/zh/posts/finops-kubernetes/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>在 Kubernetes 环境中实施 FinOps 实践，通过资源优化和成本可视化降低云基础设施支出</description><category>技术管理</category><category>FinOps</category><category>Kubernetes</category><category>成本优化</category></item><item><title>PostgreSQL 高可用架构设计与实践</title><link>https://sreai.net/zh/posts/postgresql-ha/</link><guid isPermaLink="true">https://sreai.net/zh/posts/postgresql-ha/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>设计 PostgreSQL 高可用架构，涵盖流复制、故障转移、负载均衡和备份恢复策略</description><category>编程开发</category><category>PostgreSQL</category><category>高可用</category><category>数据库</category></item><item><title>CI/CD 流水线安全左移实践</title><link>https://sreai.net/zh/posts/ci-cd-security/</link><guid isPermaLink="true">https://sreai.net/zh/posts/ci-cd-security/</guid><pubDate>Sun, 26 Jul 2026 00:00:00 +0000</pubDate><description>将安全左移到 CI/CD 流水线早期阶段，通过自动化扫描和策略即代码降低安全风险</description><category>技术管理</category><category>CI/CD</category><category>安全</category><category>DevSecOps</category></item></channel></rss>