深色模式
基础设施即代码上云
摘要:IaC 把「云资源」变成可评审、可回滚的代码。本文用 Terraform 演示一套最小可用的上云工程:目录结构、远程 state、init/plan/apply 标准流程,以及为什么绝对不能手工改动 Terraform 管理的资源。
适用环境
bash
# 安装 Terraform
which terraform || (sudo apt-get update && sudo apt-get install -y gnupg software-properties-common \
&& curl -fsSL https://apt.releases.hashicorp.com/gpg | sudo gpg --dearmor -o /usr/share/keyrings/hashicorp-archive-keyring.gpg \
&& echo "deb [signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list \
&& sudo apt-get update && sudo apt-get install -y terraform)
terraform version1
2
3
4
5
6
2
3
4
5
6
操作步骤
一、目录结构:按环境隔离
text
infra/
├── modules/
│ └── web/ # 可复用模块:输入变量 → 输出
│ ├── main.tf
│ ├── variables.tf
│ └── outputs.tf
├── envs/
│ ├── prod/
│ │ ├── main.tf
│ │ ├── backend.tf
│ │ └── terraform.tfvars
│ └── staging/
│ └── ...1
2
3
4
5
6
7
8
9
10
11
12
13
2
3
4
5
6
7
8
9
10
11
12
13
每个环境一个 state,互不影响。这是最基本也最重要的隔离手段。
二、远程 state(多人协作的前提)
hcl
# envs/prod/backend.tf
terraform {
required_version = ">= 1.5"
backend "s3" {
bucket = "my-tfstate-bucket"
key = "prod/terraform.tfstate"
region = "ap-east-1"
encrypt = true
dynamodb_table = "tf-lock" # 加锁,防止并发 apply
}
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
}1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
危险
terraform.tfstate 里包含资源的完整描述,可能含密码、私钥等敏感信息。绝不能提交到 Git,必须放远程后端并开启加密与访问权限控制。
三、写一个可复用模块
hcl
# modules/web/variables.tf
variable "name" { type = string }
variable "instance_type" { type = string, default = "t3.small" }
variable "subnet_ids" { type = list(string) }
variable "sg_ids" { type = list(string) }
# modules/web/main.tf
resource "aws_instance" "web" {
count = 2
ami = "ami-0abc123"
instance_type = var.instance_type
subnet_id = var.subnet_ids[count.index]
vpc_security_group_ids = var.sg_ids
tags = {
Name = "${var.name}-${count.index}"
env = var.name
}
}
# modules/web/outputs.tf
output "ids" { value = aws_instance.web[*].id }1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
四、标准变更流程
bash
cd envs/prod
# 1) 初始化(下载 provider、连接后端)
terraform init
# 2) 格式化与语法检查
terraform fmt -recursive
terraform validate
# 3) 预览变更 —— 这一步必须看
terraform plan -out=tfplan
# 4) 确认无误后执行
terraform apply tfplan
# 5) 查看产出
terraform output1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
注意
永远先 plan 再 apply,并且用 -out 保存 plan 文件后再 apply,确保执行的正是你审阅过的那份变更。CI 里也应把 plan 结果作为评审附件,而不是直接自动 apply。
五、变更评审要点
看 plan 输出时重点确认这三件事:
- 有没有意外的删除:
Plan: 0 to add, 0 to change, 3 to destroy这种要立刻停下来看。 - 是不是 in-place 修改还是重建(forces replacement):重建意味着服务中断。
- 含敏感值的字段:如数据库密码变更。
bash
# 只看将要发生的变化摘要
terraform plan | grep -E "^(Plan| # )" | head -30
# 阻止对某资源被误删
terraform state list | grep aws_instance1
2
3
4
2
3
4
六、导入已有资源(存量上 IaC)
bash
# 把已有实例纳入管理(先写代码,再 import)
terraform import aws_instance.web i-0abc123
# 导入后跑 plan,根据差异补全代码里的参数,直到 plan 显示 No changes
terraform plan1
2
3
4
2
3
4
验证
- [ ]
terraform validate通过,terraform fmt -check无输出 - [ ] state 存储在远程后端,本地没有
terraform.tfstate被提交 - [ ]
terraform plan在无变更时显示 "No changes" - [ ] 团队成员并发执行
apply时会被锁阻塞 - [ ] 敏感变量通过环境变量或密钥管理服务传入,未出现在代码与 state 明文里
常见坑
- 手工改了被管理的资源:下次 apply 会被「纠正」,且容易产生意外重建。改资源必须走代码。
- 本地 state 丢失:state 没了等于 Terraform 认为资源不存在,会重复创建一遍,账单翻倍。
- 版本不固定:provider 写
"~> 5.0"或不写版本,某天升级后 plan 出现大量意外变更。 - 把密码写进 tfvars 并提交:应改用变量 + 密钥管理服务/环境变量。
- 一次 apply 改太多:变更粒度过大时出问题难以回滚,建议小步提交。