深色模式
Nginx 日志与监控
摘要:想知道 Nginx 扛了多少流量、有多少 5xx、上游慢不慢,靠的是访问日志 + 状态页两件事。本文教你把访问日志改成结构化 JSON、开启
stub_status与状态页,并用 exporter 把数据送进 Prometheus。
适用环境
bash
nginx -V 2>&1 | grep -o 'with-http_stub_status_module' # 必须带此模块
sudo nginx -T | grep -n 'access_log\|error_log'
ls /var/log/nginx/1
2
3
2
3
操作步骤
1. 定义 JSON 访问日志格式
nginx
# /etc/nginx/conf.d/logformat.conf
log_format json_combined escape=json
'{'
'"time":"$time_iso8601",'
'"remote_addr":"$remote_addr",'
'"request":"$request",'
'"status":$status,'
'"body_bytes":$body_bytes_sent,'
'"request_time":$request_time,'
'"upstream_addr":"$upstream_addr",'
'"upstream_time":"$upstream_response_time",'
'"upstream_status":"$upstream_status",'
'"ua":"$http_user_agent"'
'}';
access_log /var/log/nginx/access.log json_combined;
error_log /var/log/nginx/error.log warn;1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
2. 记录 upstream 耗时(定位"慢在 Nginx 还是慢在后端")
nginx
log_format timing '$remote_addr [$time_iso8601] "$request" '
'status=$status rt=$request_time '
'uct="$upstream_connect_time" uht="$upstream_header_time" '
'urt="$upstream_response_time"';
access_log /var/log/nginx/timing.log timing;1
2
3
4
5
2
3
4
5
字段含义:uct 建连耗时、uht 收到响应头耗时、urt 完整响应耗时。三者差值能立刻区分是网络、后端排队还是后端处理慢。
3. 开启 stub_status 状态页
nginx
server {
listen 127.0.0.1:8088; # 只监听本机,禁止公网访问
location = /nginx_status {
stub_status;
access_log off;
allow 127.0.0.1;
deny all;
}
}1
2
3
4
5
6
7
8
9
2
3
4
5
6
7
8
9
bash
sudo nginx -t && sudo systemctl reload nginx
curl -s http://127.0.0.1:8088/nginx_status1
2
2
输出三行:Active connections、server accepts handled requests(三个累计计数)、Reading/Writing/Waiting。
4. 部署 nginx-prometheus-exporter
bash
# 到项目 Releases 页核对最新版本号后下载,切勿照抄过期版本
VERSION=1.1.0
curl -LO https://github.com/nginxinc/nginx-prometheus-exporter/releases/download/v${VERSION}/nginx-prometheus-exporter_${VERSION}_linux_amd64.tar.gz
tar xzf nginx-prometheus-exporter_${VERSION}_linux_amd64.tar.gz
sudo install -m 0755 nginx-prometheus-exporter /usr/local/bin/1
2
3
4
5
2
3
4
5
bash
/usr/local/bin/nginx-prometheus-exporter \
--nginx.scrape-uri=http://127.0.0.1:8088/nginx_status \
--web.listen-address=:91131
2
3
2
3
5. 让 Prometheus 抓取
yaml
# prometheus.yml 片段
scrape_configs:
- job_name: nginx
static_configs:
- targets: ['127.0.0.1:9113']1
2
3
4
5
2
3
4
5
6. 日志轮转
bash
sudo tee /etc/logrotate.d/nginx <<'EOF'
/var/log/nginx/*.log {
daily
rotate 14
missingok
compress
delaycompress
notifempty
sharedscripts
postrotate
[ -f /run/nginx.pid ] && kill -USR1 $(cat /run/nginx.pid)
endscript
}
EOF
sudo logrotate -d /etc/logrotate.d/nginx # 调试模式,先看会不会报错1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
2
3
4
5
6
7
8
9
10
11
12
13
14
15
DANGER
日志轮转后必须给 Nginx 主进程发 USR1 信号重新打开日志文件,否则进程会继续往已被重命名的旧文件里写,磁盘悄悄被撑满。
验证
bash
curl -s http://127.0.0.1:8088/nginx_status | head -3
curl -s http://127.0.0.1:9113/metrics | grep nginx_connections_active
# 统计近一分钟 5xx 数量
sudo awk -F'"' '$0 ~ /"status":5[0-9][0-9]/' /var/log/nginx/access.log | wc -l
# 用 jq 看 JSON 日志(需已安装 jq)
sudo tail -n 1 /var/log/nginx/access.log | jq .1
2
3
4
5
6
7
8
2
3
4
5
6
7
8
常见坑
WARNING
stub_status 只暴露活跃连接和累计请求数,没有状态码分布和延迟。要做"错误率""P99 延迟"必须依赖访问日志(如 mtail、filebeat 或 Loki 解析)。
WARNING
escape=json 在低版本 Nginx 上不支持,配置会报 unknown directive。可升级 Nginx,或退化为普通日志格式再交给采集端解析。
DANGER
stub_status 暴露的连接数属于敏感信息,务必用 allow 127.0.0.1; deny all; 限制来源,不要直接暴露在公网 80 端口上。