xlabs-club/awesome-x-ops
GitHub: xlabs-club/awesome-x-ops
一份系统梳理现代 X-Ops 各分支(AI Ops、LLM 可观测性、平台工程、GitOps、FinOps、DevSecOps、DataOps)生产级开源工具的精选指南,帮助团队在构建和运维生产系统时高效完成技术选型。
Stars: 20 | Forks: 6
# awesome-x-ops
一份精选的现代 X-Ops 指南:AI Ops、LLM/Agent Observability、Platform Engineering、GitOps、DataOps、FinOps、DevSecOps 以及生产级开源运维工具。
语言:English | [简体中文](README.zh-CN.md)
[](https://awesome.re)
[](CONTRIBUTING.md)
[](https://creativecommons.org/licenses/by-nc/4.0/)
## 为什么选择 awesome-x-ops?
运维工作不再仅仅是基础设施监控或 CI/CD 粘合剂。现代团队需要一份涵盖 AI 原生应用、LLM observability、Platform Engineering、软件交付、云成本、安全性和开发者体验的实用指南。
本列表专注于帮助团队构建、运行、观察、保护和优化生产系统的工具。
## 特色指南
- [LLM 和 Agent Observability 技术栈](#llm-and-agent-observability):用于 LLM 和 agent 系统的 tracing、prompt 监控、评估、反馈和生产遥测。
- [AI 基础设施](#ai-infrastructure):Web 爬取、适配 AI 的提取、搜索智能和 RAG 数据获取。
- [Platform Engineering 技术栈](#platform-engineering):内部开发者平台、IAM、IaC、artifacts、API 工具、CI/CD 和测试。
- [GitOps 和 Kubernetes 运维技术栈](#kubernetes-operations):集群网络、autoscaling、部署和 runtime 运维。
- [FinOps 技术栈](#finops):云和 Kubernetes 成本可视化、分配和预测。
- [DevSecOps 和软件供应链技术栈](#security-and-supply-chain):策略、runtime 安全、SBOM、扫描和软件供应链风险管理。
- [DataOps 技术栈](#dataops):数据流、编排和数据资产生命周期工具。
## 目标读者
- 构建内部开发者平台的 Platform Engineering 团队。
- 正在实现运维技术栈现代化的 DevOps、SRE 和基础设施团队。
- 在生产环境中运行 LLM、RAG 和 agent 应用的 AI 工程团队。
- 在购买或构建之前寻找可靠开源方案的工程领导者。
- 希望其生产级运维工具能够被发现的 open-source 维护者。
## 筛选原则
- 保持条目简洁、高效、准确且相关。
- 如果存在可靠的项目仓库,优先选择 GitHub 链接。
- 仅包含经过验证、可靠且高质量的项目。
- 忽略重复项或已被等效条目涵盖的项目。
- 在有用时添加或完善分类,但避免不相关的内容。
- 优先选择生产级 open source,而不是演示、废弃的实验或仅限供应商的营销页面。
## LLM 和 Agent Observability
用于在生产环境中对 LLM、RAG 和 agent 应用进行 tracing、评估、调试和运维的工具。
- [LiteLLM](https://github.com/BerriAI/litellm):兼容 OpenAI 的 LLM gateway,具备路由、预算、日志记录和 provider 抽象功能。
- [Langfuse](https://github.com/langfuse/langfuse):开源的 LLM 工程平台,用于 traces、prompt 管理、评估和指标。
- [DeepEval](https://github.com/confident-ai/deepeval):用于在 CI 或生产工作流中测试 RAG、agent 和模型输出的 LLM 评估框架。
- [Ragas](https://github.com/explodinggradients/ragas):用于 RAG pipeline 和 LLM 应用的评估框架。
- [Arize Phoenix](https://github.com/Arize-ai/phoenix):用于 LLM、RAG 和 ML 系统的开源 observability 和评估平台。
- [OpenInference](https://github.com/Arize-ai/openinference):用于 tracing LLM、RAG 和 agent 应用的 OpenTelemetry instrumentation 和语义约定。
- [OpenLLMetry](https://github.com/traceloop/openllmetry):基于 OpenTelemetry 的 LLM 应用和 agent 工作流 observability 工具。
- [Helicone](https://github.com/Helicone/helicone):用于监控 LLM 使用情况、延迟、成本、缓存和请求日志的开源 observability 平台。
- [OpenLIT](https://github.com/openlit/openlit):原生支持 OpenTelemetry 的 AI 工程平台,用于 LLM observability、评估、guardrails、prompt 管理和 GPU 监控。
- [LangWatch](https://github.com/langwatch/langwatch):用于 LLM 监控、评估、traces 和 agent 测试的开源平台。
- [Opik](https://github.com/comet-ml/opik):用于 tracing、评估和监控 LLM 应用、RAG 系统和 agent 工作流的开源平台。
- [promptfoo](https://github.com/promptfoo/promptfoo):用于 prompt 测试、LLM 评估、红蓝对抗和 CI/CD 回归检查的开源 CLI 和平台。
- [Langtrace](https://github.com/Scale3-Labs/langtrace):基于 OpenTelemetry 的 observability 平台,用于 tracing、评估和监控 LLM 应用。
- [Future AGI](https://github.com/future-agi/future-agi):可自托管的平台,用于评估、观察和改进 LLM 及 AI agent 应用。
- [CozeLoop](https://github.com/coze-dev/coze-loop):AI agent 优化平台,涵盖开发、调试、评估和生产监控工作流。
- [Agenta](https://github.com/Agenta-AI/agenta):开源的 LLMOps 平台,用于 prompt 管理、playgrounds、评估和 observability。
- [abtop](https://github.com/graykode/abtop):htop 风格的终端监视器,用于 AI 编程 agent 会话、token、上下文窗口、速率限制和端口。
- [agenttrace](https://github.com/luoyuctl/agenttrace):本地优先的 TUI,用于检查 AI 编程 agent 的成本、token、延迟、故障和报告。
- [ax](https://github.com/Necmttn/ax):用于 AI 编程 agent 的本地优先遥测和内存图,涵盖成本、工具、技能、会话和 OTLP 事件。
- [TensorZero](https://github.com/tensorzero/tensorzero):开源的 LLMOps 平台,结合了 LLM gateway、observability、评估、优化和实验。
- [Evidently](https://github.com/evidentlyai/evidently):开源的 ML 和 LLM observability 框架,用于评估、测试、监控和数据质量检查。
- [RagaAI Catalyst](https://github.com/raga-ai-hub/RagaAI-Catalyst):用于 tracing、调试和监控多 agent LLM 系统的 Agent AI observability 和评估 SDK。
- [Pydantic Logfire](https://github.com/pydantic/logfire):用于 tracing 和监控生产级 LLM 及 agent 系统的 AI observability 平台。
- [Laminar](https://github.com/lmnr-ai/lmnr):专为 AI agent 和 LLM 应用构建的开源 observability 平台。
- [MLflow](https://github.com/mlflow/mlflow):开源的 AI 工程平台,用于调试、评估、监控和优化 agent、LLM 和 ML 模型。
- [Giskard](https://github.com/Giskard-AI/giskard-oss):用于 LLM 应用和 AI agent 的开源评估和测试框架。
- [ZenML](https://github.com/zenml-io/zenml):用于生产级 ML、LLM 和 agent pipeline 的 AI 平台,具备编排、追踪和部署工作流。
- [Guardrails](https://github.com/guardrails-ai/guardrails):用于验证 LLM 输出并强制执行安全性、质量和结构化响应检查的框架。
- [Plano](https://github.com/katanemo/plano):面向 agentic 应用的 AI 原生代理和数据平面,具备路由、安全、编排和 observability 功能。
- [AgentSight](https://github.com/eunomia-bpf/agentsight):基于 eBPF 的系统级 tracing,无需应用 instrumentation 即可观察 AI agent 执行。
- [AgentOps](https://github.com/AgentOps-AI/agentops):Python SDK,用于监控 AI agent、追踪 LLM 成本、对运行进行基准测试,并与常见的 agent 框架集成。
- [Portkey AI Gateway](https://github.com/Portkey-AI/gateway):AI gateway,用于路由 LLM 流量、应用 guardrails 并集中管理生产应用的模型访问。
- [BISHENG](https://github.com/dataelement/bisheng):面向企业 AI 应用的开放 LLM DevOps 平台,具备 GenAI 工作流、RAG、agent、模型管理、评估、数据集和 observability 功能。
- [OpenObserve](https://github.com/openobserve/openobserve):开源的 observability 平台,用于日志、指标、traces、前端监控、pipeline 和 LLM observability。
- [MCP Gateway](https://github.com/IBM/mcp-context-forge):用于 MCP、A2A 和 API 工具的 AI gateway、注册中心和代理,具备集中发现、guardrails 和管理功能。
- [NVIDIA NeMo Guardrails](https://github.com/NVIDIA-NeMo/Guardrails):用于向基于 LLM 的对话系统添加可编程安全、对话和策略 guardrails 的工具包。
- [Llama Guard](https://github.com/meta-llama/PurpleLlama):Meta 的开放信任与安全工具包,用于评估和过滤 LLM 的输入、输出及模型风险。
- [LLM Guard](https://github.com/protectai/llm-guard):安全工具包,用于净化 LLM 的输入和输出、检测 prompt 注入、拦截有害内容并减少数据泄露。
- [OpenEvals](https://github.com/langchain-ai/openevals):现成的评估器,用于在开发和 CI 工作流中对 LLM 应用进行测试和回归检查。
- [Envoy AI Gateway](https://github.com/envoyproxy/ai-gateway):基于 Envoy 的 gateway,用于管理跨 provider 和平台对生成式 AI 服务的统一访问。
- [Inference Gateway](https://github.com/inference-gateway/inference-gateway):云原生 LLM gateway,用于统一 provider、路由推理流量,并在 Kubernetes 上暴露对 OpenTelemetry 友好的运维操作。
- [OneAIFW](https://github.com/funstory-ai/aifw):轻量级本地 AI 防火墙,用于在 LLM 调用前对敏感数据进行匿名化,并在响应后恢复。
- [Microsoft MCP Gateway](https://github.com/microsoft/mcp-gateway):反向代理和管理层,用于运维 MCP server,具备会话感知路由和 Kubernetes 生命周期支持。
- [CoAI](https://github.com/coaidev/coai):多租户 AI 平台,具备统一的 LLM gateway、provider 路由、成本管理、计费和模型缓存,适用于企业部署。
## AI 服务和推理运维
用于在生产环境中部署、扩展、路由和运维 AI 模型推理工作负载的工具。
- [Ray Serve](https://github.com/ray-project/ray):Ray 中的可扩展模型服务库,用于构建分布式在线推理 API 和 LLM 服务工作负载。
- [Triton Inference Server](https://github.com/triton-inference-server/server):优化的推理 server,用于跨 GPU、CPU 以及云或边缘环境部署 AI 模型。
- [KServe](https://github.com/kserve/kserve):Kubernetes 原生平台,用于标准化、可扩展的生成式和预测式 AI 推理服务。
- [AIBrix](https://github.com/vllm-project/aibrix):云原生基础设施组件,用于成本高效、可扩展的 GenAI 和 LLM 推理运维。
- [Dynamo](https://github.com/ai-dynamo/dynamo):用于数据中心级 LLM 和生成式 AI 工作流的分布式推理服务框架,具备面向 Kubernetes 的路由和扩展功能。
- [dstack](https://github.com/dstackai/dstack):与供应商无关的控制平面,用于跨云、Kubernetes 和裸金属环境配置 GPU 并编排训练、推理和 agent 工作负载。
- [llm-d](https://github.com/llm-d/llm-d):Kubernetes 原生分布式推理技术栈,用于在现代加速器上通过智能路由实现高性能 LLM 服务。
- [KubeAI](https://github.com/kubeai-project/kubeai):Kubernetes AI 推理 operator,用于通过兼容 OpenAI 的 API 服务 LLM、VLM、embedding 和语音模型。
## AIOps
- [Netdata](https://github.com/netdata/netdata):分布式实时监控,用于基础设施指标、可视化和告警。
- [Apache HertzBeat](https://github.com/apache/hertzbeat):Apache 实时 observability 和监控系统,具备无 agent 采集、告警、状态页面和 AI 辅助运维功能。
- [PostHog](https://github.com/PostHog/posthog):开源的产品分析平台,用于用户行为追踪和产品指标。
- [SREWorks](https://github.com/alibaba/SREWorks):云原生 DataOps 和 AIOps 平台,用于运维基于 Kubernetes 的应用和基础设施。
## AI 基础设施
用于 Web 爬取、适配 AI 的提取、搜索智能和 RAG 数据获取工作流的基础设施。
- [Firecrawl](https://github.com/firecrawl/firecrawl):Web 搜索、抓取、爬取和提取 API,可将 Web 数据转换为适配 LLM 的 Markdown 和结构化输出。
- [Crawl4AI](https://github.com/unclecode/crawl4ai):开源的适配 LLM 的 Web 爬虫和抓取工具,用于构建 RAG、agent 和 Web 数据 pipeline。
- [Open SEO](https://github.com/every-app/open-seo):开源的 SEO 和搜索智能平台,用于关键词研究、站点审计、反向链接分析以及 Google Search Console MCP 工作流。
## Agentic 工作流
- [AutoGPT](https://github.com/Significant-Gravitas/Auto-GPT):能够分解并执行复杂任务的自主 AI agent 框架。
- [Langflow](https://github.com/langflow-ai/langflow):用于 LangChain 风格 L 工作流的图形化构建器。
- [Dify](https://github.com/langgenius/dify):开源的 LLM 应用开发平台,具备可视化的 agent 工作流和 AI 应用部署。
- [LangChain](https://github.com/langchain-ai/langchain):用于构建 LLM 驱动应用的框架,包括 agent 工作流编排。
- [Flowise](https://github.com/FlowiseAI/Flowise):低代码 LLM 工作流编排工具,用于可视化构建 AI 应用链。
- [crewAI](https://github.com/crewAIInc/crewAI):用于协作型 AI agent 的框架,具备角色定义和任务编排功能。
- [LlamaIndex](https://github.com/run-llama/llama_index):用于 LLM 应用的数据框架,支持结构化数据检索和增强。
- [Haystack](https://github.com/deepset-ai/haystack):可扩展的框架,用于问答和自定义 AI 工作流开发。
- [BentoML](https://github.com/bentoml/BentoML):开源的模型服务平台,用于跨框架部署模型并编排 AI 应用。
- [trpc-agent-go](https://github.com/trpc-group/trpc-agent-go):用于生产级 agent 系统的 Go 框架,具备图工作流、工具、记忆、评估和 observability。
## DataOps
- [Dagster](https://dagster.io/):数据编排平台,用于对数据资产进行建模和管理数据生命周期。
- [Apache NiFi](https://nifi.apache.org/):可视化数据流编排,用于跨系统路由、转换和协调数据。
- [DataHub](https://github.com/datahub-project/datahub):元数据平台,用于现代数据和 AI 技术栈中的数据发现、血缘、治理和 observability。
- [OpenMetadata](https://github.com/open-metadata/OpenMetadata):统一的元数据平台,用于数据发现、血缘、治理和数据 observability。
- [Great Expectations](https://github.com/great-expectations/great_expectations):数据质量框架,用于验证数据集、记录预期并捕获 pipeline 回归。
- [Soda Core](https://github.com/sodadata/soda-core):数据契约和质量检查引擎,用于验证现代数据技术栈中的数据 pipeline。
- [Elementary](https://github.com/elementary-data/elementary):dbt 原生的数据 observability 平台,用于监控 pipeline、测试、新鲜度和异常。
- [OpenLineage](https://github.com/OpenLineage/OpenLineage):用于跨数据 pipeline 和平台收集血缘元数据的开放标准和工具。
- [Marquez](https://github.com/MarquezProject/marquez):元数据服务,用于跨作业和数据集收集、聚合和可视化数据血缘。
- [Temporal](https://github.com/temporalio/temporal):持久执行平台,用于构建可靠的工作流、后台作业和长时间运行的业务流程。
- [Kestra](https://github.com/kestra-io/kestra):事件驱动的编排和调度平台,用于声明式数据、基础设施和运维工作流。
### 流式运维
- [Kafbat UI](https://github.com/kafbat/kafka-ui):开源 Web UI,用于管理 Apache Kafka 集群、topic、消费者、schema 和 Kafka Connect。
- [Apache SeaTunnel](https://github.com/apache/seatunnel):分布式数据集成平台,用于海量批处理和流式数据迁移。
## FinOps
- [Infracost](https://github.com/infracost/infracost):云成本预测工具,用于 Terraform 和 Kubernetes 成本估算。
- [kubecost](https://kubecost.com/):Kubernetes 成本管理和监控平台。
- [OpenCost](https://opencost.io/):开源工具,用于在 Kubernetes 环境中追踪和分配云成本。
- [OptScale](https://github.com/hystax/optscale):开源的 FinOps 和云成本优化平台,适用于 AWS、Azure、GCP、Alibaba Cloud 和 Kubernetes。
- [Cloud Custodian](https://github.com/cloud-custodian/cloud-custodian):策略即代码规则引擎,用于云治理、成本优化和自动化资源操作。
- [KubeStellar Console](https://github.com/kubestellar/console):多集群 Kubernetes 仪表盘,具备 AI 驱动的运维、实时 observability 以及跨边缘和云集群的 CNCF 项目集成。
## 可观测性
- [Prometheus](https://github.com/prometheus/prometheus):监控系统的时间序列数据库,广泛用于云原生指标和告警。
- [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics):快速、高性价比的时间序列数据库和监控技术栈,用于大规模的 Prometheus 兼容指标。
- [Grafana Mimir](https://github.com/grafana/mimir):具备水平扩展能力、多租户的长期存储后端,用于 Prometheus 指标。
- [Grafana Tempo](https://github.com/grafana/tempo):分布式 tracing 后端,用于海量 trace 存储,索引开销极低。
- [Perses](https://github.com/perses/perses):CNCF observability 可视化项目,用于跨 Prometheus、Tempo、Loki 和相关数据源构建仪表盘。
- [Grafana Loki](https://github.com/grafana/loki):日志聚合系统,旨在高效索引标签并与 Grafana 集成。
- [OpenTelemetry Collector](https://github.com/open-telemetry/opentelemetry-collector):供应商中立的收集器,用于接收、处理和导出遥测数据。
- [SigNoz](https://github.com/SigNoz/signoz):OpenTelemetry 原生的 observability 平台,集成了指标、traces、日志、仪表盘和告警。
- [Jaeger](https://github.com/jaegertracing/jaeger):CNCF 分布式 tracing 平台,用于监控和排查 microservice。
- [Vector](https://github.com/vectordotdev/vector):高性能 observability 数据 pipeline,用于收集、转换和路由日志和指标。
- [Grafana Alloy](https://github.com/grafana/alloy):带有可编程 pipeline 的 OpenTelemetry Collector 发行版,用于收集、处理和转发 observability 信号。
- [Pixie](https://github.com/pixie-io/pixie):Kubernetes 原生的 observability 平台,使用 eBPF 捕获指标、事件、traces 和网络遥测,无需手动 instrumentation。
- [Parca](https://github.com/parca-dev/parca):持续分析平台,用于分析 CPU 和内存使用情况以提高性能、可靠性和基础设施效率。
- [Kepler](https://github.com/sustainable-computing-io/kepler):Kubernetes 功耗和能耗 exporter,用于配合 Prometheus 测量容器、pod 和节点的能耗。
- [Inspektor Gadget](https://github.com/inspektor-gadget/inspektor-gadget):基于 eBPF 的检查工具包,用于收集底层的 Kubernetes 和 Linux 运维遥测数据。
- [Robusta](https://github.com/robusta-dev/robusta):Kubernetes 告警丰富化和自动化平台,用于处理 Prometheus 告警、runbook 和修复工作流。
- [Coroot](https://github.com/coroot/coroot):开源的 observability 和 APM 平台,具备指标、日志、traces、性能分析、SLO 和 AI 辅助根因分析。
## Kubernetes 运维
- [Cilium](https://github.com/cilium/cilium):基于 eBPF 的 Kubernetes 网络、安全和 observability 平台。
- [Headlamp](https://github.com/kubernetes-sigs/headlamp):可扩展的 Kubernetes Web UI,用于集群可视化、资源管理和运维插件。
- [cert-manager](https://github.com/cert-manager/cert-manager):Kubernetes 原生的证书管理 controller,用于签发和续订 TLS 证书。
- [KEDA](https://github.com/kedacore/keda):Kubernetes 事件驱动的 autoscaler,用于根据外部指标和事件源扩展工作负载。
- [Velero](https://github.com/velero-io/velero):Kubernetes 备份、恢复和迁移工具,用于集群资源和持久卷。
- [External Secrets Operator](https://github.com/external-secrets/external-secrets):Kubernetes operator,将来自外部密钥管理器的 secret 同步到 Kubernetes Secret 中。
- [Reloader](https://github.com/stakater/Reloader):Kubernetes controller,当引用的 ConfigMap 或 Secret 发生变化时触发工作负载的滚动重启。
- [Karpenter](https://github.com/kubernetes-sigs/karpenter):灵活的 Kubernetes 节点 autoscaler,用于提高集群效率和工作负载调度。
- [Koordinator](https://github.com/koordinator-sh/koordinator):Kubernetes 调度系统,用于工作负载混部、资源优化和成本感知的集群运维。
- [Capsule](https://github.com/projectcapsule/capsule):Kubernetes 多租户框架,允许平台团队通过基于策略的租户边界委派 namespace。
- [vCluster](https://github.com/loft-sh/vcluster):在 namespace 内运行的虚拟 Kubernetes 集群,用于多租户、隔离和 Platform Engineering 工作流。
- [Chaos Mesh](https://github.com/chaos-mesh/chaos-mesh):Kubernetes 原生的混沌工程平台,用于在受控故障下测试系统弹性。
- [Goldilocks](https://github.com/FairwindsOps/goldilocks):Kubernetes 资源建议仪表盘,帮助根据 VPA 洞察调整工作负载的 request 和 limit。
- [Glasskube](https://github.com/glasskube/glasskube):Kubernetes 包管理器,提供 GUI 和 CLI 支持,适用于具备依赖感知、GitOps 就绪的应用运维。
- [Botkube](https://github.com/kubeshop/botkube):Kubernetes ChatOps 助手,用于监控集群、展示事件并帮助团队调试部署。
- [mirrord](https://github.com/metalbear-co/mirrord):Kubernetes 开发工具,允许本地进程在集群的网络、环境和流量上下文中运行。
- [OpenKruise](https://github.com/openkruise/kruise):CNCF Kubernetes 工作负载自动化套件,用于高级的应用部署、扩展和生命周期管理。
- [kOps](https://github.com/kubernetes/kops):生产级 Kubernetes 集群生命周期工具,用于跨云环境的安装、升级和运维。
- [KubeOne](https://github.com/kubermatic/kubeone):Kubernetes 集群生命周期管理工具,用于跨云、本地、边缘和 IoT 环境自动化运维。
## 安全和供应链
- [Falco](https://github.com/falcosecurity/falco):CNCF runtime 安全工具,用于检测容器和 Kubernetes 中的可疑行为。
- [Kyverno](https://github.com/kyverno/kyverno):Kubernetes 原生的策略引擎,用于验证、变更、生成和镜像验证。
- [Open Policy Agent](https://github.com/open-policy-agent/opa):通用策略引擎,支持跨 Kubernetes、CI/CD、API 和基础设施的策略即代码。
- [Gatekeeper](https://github.com/open-policy-agent/gatekeeper):Kubernetes 准入 controller,用于跨集群强制执行 OPA 策略和审计约束。
- [Syft](https://github.com/anchore/syft):用于从容器镜像和文件系统生成 SBOM 的 CLI 和库。
- [Grype](https://github.com/anchore/grype):用于容器镜像和文件系统的漏洞扫描器,与 Syft 生成的 SBOM 配合良好。
- [Kubescape](https://github.com/kubescape/kubescape):Kubernetes 安全平台,用于风险分析、合规性、配置错误扫描以及 CI/CD 或集群检查。
- [Gitleaks](https://github.com/gitleaks/gitleaks):密钥扫描器,用于检测 Git 仓库、文件和 CI/CD 工作流中硬编码的凭证。
- [TruffleHog](https://github.com/trufflesecurity/trufflehog):密钥扫描器,用于查找、验证和分析跨 Git、文件系统、CI 日志和云来源中泄露的凭证。
- [Prowler](https://github.com/prowler-cloud/prowler):多云安全和合规平台,用于审计 AWS、Azure、GCP、Kubernetes 和 SaaS 环境。
- [KubeArmor](https://github.com/kubearmor/KubeArmor):Kubernetes runtime 安全强制执行系统,通过基于 LSM 的策略实现最小权限工作负载强化。
- [Kubewarden](https://github.com/kubewarden/adm-controller):Kubernetes 准入策略引擎,运行 WebAssembly 策略以实现策略即代码治理。
- [cosign](https://github.com/sigstore/cosign):Sigstore 工具,用于签名和验证容器镜像、blob 和软件制品,并支持透明日志。
- [SLSA GitHub Generator](https://github.com/slsa-framework/slsa-github-generator):GitHub Actions 工作流,用于为构建和发布制品生成 SLSA 来源证明。
- [Chainloop](https://github.com/chainloop-dev/chainloop):软件供应链控制平面,用于收集 SDLC 证据、证明、SBOM、VEX、SARIF 和策略检查。
- [SafeDep vet](https://github.com/safedep/vet):策略即代码工具,用于检测恶意、存在漏洞或有风险的 open-source 包依赖。
- [OSV-Scanner](https://github.com/google/osv-scanner):漏洞扫描器,使用 OSV.dev 数据在源代码、lockfile、SBOM 和容器镜像中发现已知漏洞。
- [OpenSSF Scorecard](https://github.com/ossf/scorecard):open-source 项目的自动化安全健康检查器,涵盖依赖、CI/CD、分支保护和漏洞治理信号。
- [GUAC](https://github.com/guacsec/guac):软件供应链图,聚合 SBOM、SLSA 证明、漏洞和依赖元数据以进行风险分析。
- [ORT](https://github.com/oss-review-toolkit/ort):用于在依赖、许可证、版权、漏洞和 SBOM 生成方面自动化 open-source 合规性检查的工具包。
- [CycloneDX CLI](https://github.com/CycloneDX/cyclonedx-cli):命令行工具,用于验证、转换、合并和比对 CycloneDX SBOM 及相关格式。
- [Dependency-Track](https://github.com/DependencyTrack/dependency-track):组件分析平台,用于追踪 SBOM、漏洞、许可证和软件供应链风险。
## Platform Engineering
为 Platform Engineering 精选的技术栈和工具链。
### API 管理工具
- [Hoppscotch](https://github.com/hoppscotch/hoppscotch):适用于 REST、GraphQL 和 WebSocket 的轻量级 API 开发套件。
- [Bruno](https://github.com/usebruno/bruno):快速、对 Git友好的开源 API client,用于通过桌面应用或 CLI 管理 API 集合并发起 API 调用。
### Artifact 管理
- [Harbor](https://github.com/goharbor/harbor):具备安全扫描和访问控制的企业级容器镜像仓库。
- [Skopeo](https://github.com/containers/skopeo):用于检查、复制和签名容器镜像的开源工具。
- [Nexus Repository](https://github.com/sonatype/nexus-public):通用 artifact 仓库,支持 Maven、npm、Docker 等。
- [ORAS](https://github.com/oras-project/oras):用于将任意内容存储为 OCI artifact 的工具。
### CI/CD
- [Apache Airflow](https://airflow.apache.org/):用于数据 pipeline 的开源工作流编排平台。
- [Harness](https://github.com/harness/harness):开源的端到端开发者平台,提供源代码控制、CI/CD pipeline、托管开发环境和 artifact 仓库。
- [Jenkins](https://www.jenkins.io/):开源的 CI/CD 自动化 server,拥有庞大的插件生态。
- [argo-cd](https://argo-cd.readthedocs.io/):适用于 Kubernetes 的流行声明式 GitOps CD 工具。
- [Argo Rollouts](https://github.com/argoproj/argo-rollouts):Kubernetes 渐进式交付 controller,支持蓝绿部署、金丝雀部署和基于实验的部署。
- [argo-workflows](https://github.com/argoproj/argo-workflows):Kubernetes 原生的工作流引擎。
- [Tekton](https://tekton.dev/):Kubernetes 原生的 CI/CD 框架,具备灵活的任务编排。
- [Flux](https://fluxcd.io/):流行的 Kubernetes GitOps 工具包。
- [PipeCD](https://github.com/pipe-cd/pipecd):CNCF 持续交付平台,适用于跨多个部署目标的应用、基础设施和平台运维。
### 代码服务
- [Trivy](https://github.com/aquasecurity/trivy):面向容器、代码、漏洞、配置错误和 SBOM 的综合扫描器。
- [SonarQube](https://github.com/SonarSource/sonarqube):支持 27+ 编程语言的持续代码质量平台。
- [reviewdog](https://github.com/reviewdog):面向多种语言和 linter 的自动代码审查和分析工具。
- [Dependency Track](https://dependencytrack.org/):开源的软件组件分析平台,用于供应链风险、SBOM 分析和许可证检查。
- [OpenRewrite](https://docs.openrewrite.org):大规模自动代码重构和现代化工具。
- [Hyades](https://github.com/DependencyTrack/hyades):旨在稳定后取代 Dependency-Track 的下一代软件供应链安全平台。
### Event Mesh
- [CloudEvents](https://cloudevents.io/):用于可互操作事件驱动系统的规范。
- [Argo Events](https://argoproj.github.io/argo-events/):面向 Kubernetes 的事件驱动工作流自动化框架。
- [Apache EventMesh](https://eventmesh.apache.org/):支持多种消息协议和事件流管理的分布式事件中间件。
### 功能管理和实验
- [GrowthBook](https://github.com/growthbook/growthbook):开源的功能开关、实验和产品分析平台,用于实现更安全的渐进式交付。
- [Flagsmith](https://github.com/Flagsmith/flagsmith):开源的功能开关和远程配置服务,支持自托管或托管式发布控制。
### 基础设施即代码
IaC 通过代码而非人工流程来管理和配置基础设施。
- [OpenTofu](https://github.com/opentofu/opentofu):由 Linux Foundation 管理的社区驱动的 Terraform 分支。
- [Pulumi](https://github.com/pulumi/pulumi):使用熟悉的编程语言在任何云上构建基础设施的 IaC 工具。
- [sops](https://github.com/getsops/sops):使用 AWS KMS、GCP KMS、Azure Key Vault、age 或 PGP 加密 YAML、JSON、ENV、INI 和二进制文件的编辑器。
- [Crossplane](https://github.com/crossplane/crossplane):Kubernetes 附加组件,允许平台团队组合来自多个供应商的基础设施并公开更高层次的自助服务 API。
- [Terragrunt](https://github.com/gruntwork-io/terragrunt):用于 DRY 配置和远程状态管理的 Terraform 包装器。
- [bitnami/sealed-secrets](https://github.com/bitnami-labs/sealed-secrets):通过加密 secret 以便安全存储在 Git 中并在集群内解密,实现声明式 Kubernetes secret 管理。
- [Checkov](https://github.com/bridgecrewio/checkov):用于基础设施即代码安全和合规性的静态分析工具。
- [helmfile](https://github.com/helmfile):用于编排和部署 Helm chart 的声明式工具。
- [Atlantis](https://github.com/runatlantis/atlantis):面向 Terraform 工作流、plan、apply 和协作式基础设施审查的 Pull Request 自动化工具。
### 身份和访问管理 (IAM)
信任很难。知道该信任谁更是难上加难。
- [keycloak](https://github.com/keycloak/keycloak):适用于现代应用和服务的开源 IAM。
- [OpenBao](https://github.com/openbao/openbao):开源的密钥管理系统,用于存储和分发密钥、证书和 key。
- [oauth2-proxy](https://github.com/oauth2-proxy/oauth2-proxy):适用于 Google、Azure、OpenID Connect 等的轻量级 OAuth2 反向代理,带有简单的授权检查。
- [zitadel](https://github.com/zitadel/zitadel):注重简洁性的开源 IAM,适用于现代应用和服务。
- [Casdoor](https://github.com/casdoor/casdoor):开源的身份管理平台,支持 OAuth 2.0、OIDC 和 SAML。
- [dexidp/dex](https://github.com/dexidp/dex):轻量级可插拔的 OpenID Connect (OIDC) 和 OAuth 2.0 provider。
- [pomerium](https://github.com/pomerium/pomerium):具备更丰富访问控制功能的身份感知代理。
### 内部开发者平台 (IDP)
内部开发者平台不仅仅是一堆工具的堆砌;它不是又一个管理控制台或仪表盘。
- [backstage](https://github.com/backstage/backstage):用于构建开发者门户的开放平台,帮助团队构建、部署和维护软件。
- [Kratix](https://github.com/syntasso/kratix):用于构建平台 API 的框架,允许团队在 Kubernetes 上组合和运维内部平台。
- [OpenChoreo](https://github.com/openchoreo/openchoreo):面向 Kubernetes 的开源开发者平台,具备由 Backstage 驱动的门户、CI/CD、GitOps、observability 和平台抽象。
- [KubeVela](https://github.com/kubevela/kubevela):CNCF 应用交付平台,用于在混合和多集群环境中管理 Kubernetes 工作负载。
- [Score](https://github.com/score-spec/spec):与平台无关的工作负载规范,用于一次描述服务并生成特定于环境的平台配置。
- [Superplane](https://github.com/superplanehq/superplane):面向跨服务、pipeline 和环境的 Platform Engineering 工作流的开源控制平面。
### IaaS 工具
适用于本地 Kubernetes 和容器平台调试的轻量级虚拟化工具。
- [Minikube](https://github.com/kubernetes/minikube):本地 Kubernetes 集群部署工具。
- [Vagrant](https://github.com/hashicorp/vagrant):跨平台虚拟机管理工具,支持多种虚拟化后端。
- [lima](https://github.com/lima-vm/lima):具备自动文件共享和端口转发的 Linux 虚拟机,包括异构虚拟机模拟。
- [multipass](https://github.com/canonical/multipass):Ubuntu 提供的轻量级虚拟化工具。
### 测试工具
面向测试工程师和以质量为核心的平台团队的工具。
- [googletest](https://github.com/google/googletest):Google 测试和模拟框架。
- [Selenium](https://github.com/SeleniumHQ/selenium):用于 Web 应用测试的浏览器自动化框架。
- [grafana/k6](https://github.com/grafana/k6):使用 Go 和 JavaScript 的现代化负载测试工具,也适用于 API 测试工作流。
- [JMeter](https://github.com/apache/jmeter):基于 Java 的性能测试工具,支持多种协议。
- [Tracetest](https://github.com/kubeshop/tracetest):基于 OpenTelemetry 的 trace 测试工具,用于验证分布式工作流和 observability instrumentation。
## License
本文档基于 [CC BY-NC 4.0][] 授权。
标签:DataOps, DevSecOps, Docker镜像, FinOps, GitOps, MLOps, 上游代理, 子域名突变, 平台工程, 用户代理, 运维