深色模式
Fluentd 路由与过滤
Fluentd 用统一的 JSON 事件模型,配合
tag路由和filter改写,非常适合多源聚合。本文演示采集 nginx、syslog 并统一加字段后发到 ES。
适用环境
bash
# 确认 Ruby 环境(Fluentd 基于 Ruby)
ruby -v
# 安装 td-agent(官方打包版)
curl -L https://toolbelt.treasuredata.com/sh/install-ubuntu-jammy-td-agent4.sh | sh1
2
3
4
2
3
4
操作步骤
1. 配置多源输入(/etc/td-agent/td-agent.conf)
xml
<source>
@type tail
path /var/log/nginx/access.log
tag nginx.access
<parse>
@type nginx
</parse>
</source>
<source>
@type syslog
port 5140
tag syslog.local
</source>1
2
3
4
5
6
7
8
9
10
11
12
13
14
2
3
4
5
6
7
8
9
10
11
12
13
14
2. 用 filter 统一加环境和改写字段
xml
<filter nginx.access>
@type record_transformer
<record>
env production
hostname ${hostname}
</record>
</filter>1
2
3
4
5
6
7
2
3
4
5
6
7
3. 按 tag 路由到不同输出
xml
<match nginx.access>
@type elasticsearch
host es-host
port 9200
logstash_format true
</match>
<match syslog.*>
@type stdout
</match>1
2
3
4
5
6
7
8
9
10
2
3
4
5
6
7
8
9
10
bash
sudo systemctl restart td-agent1
验证
bash
# 看 Fluentd 处理日志
sudo tail -f /var/log/td-agent/td-agent.log
# 用伪 syslog 发一条测试
logger -n localhost -P 5140 "fluentd test message"1
2
3
4
2
3
4
常见坑
WARNING
tag 是路由的唯一依据,tag 命名要有层级(如 service.env.type),避免后续 match 规则冲突。
DANGER
Fluentd 默认缓冲在内存,进程重启会丢缓冲数据。生产务必配置 buffer 用 file 缓冲并限制 chunk_limit_size。