离线流量分析的第一步通常不是马上选定一个解析库,而是先弄清楚 pcap 文件里到底留下了哪些字节。pcap 格式由全局文件头、一组报文记录头和链路层原始报文三部分组成,这种布局看似简单,实际处理时却会碰到字节序识别、字段截断、结构体对齐和协议栈偏移计算等细节。本文会以 C++ 为例,先手动解析二进制结构,再讨论 libpcap 的封装路径,最后落在一个可扩展的 TCP 五元组提取实现上。

一、Pcap 全局头:先解决字节序和版本问题
每个 pcap 文件最前面的 24 个字节是全局头。它记录了这个抓包文件使用的字节序、版本号、最大捕获长度和链路层类型。C++ 里可以通过结构体来描述这份布局,但一定要关闭编译器自动填充,或者干脆按手工偏移读取,否则 24 字节长度会不一致。
全局头中最关键的字段是 magic number,它的标准值是 0xa1b2c3d4。由于抓包工具在写文件时可能使用小端或大端格式,读取到的四个字节顺序可能正好相反。如果程序直接读取成 0xa1b2c3d4,说明文件字节序与当前主机一致;如果读到 0xd4c3b2a1,说明后续所有多字节字段都需要做字节交换。snaplen 表示每个包最多保存多少字节,network 字段则决定后面链路层数据的类型,常见的 1 表示以太网。
下面这段代码定义了全局头结构,并给出 byte swap 函数。对于需要跨平台运行的解析器来说,不建议用 reinterpret_cast 把文件字节直接当作结构体使用,更稳妥的做法是先读入缓冲,再按已知偏移逐字段解析,这样可以把字节序和填充问题完全控制在自己手里。
#include <cstdint>
#include <fstream>
#include <vector>
#include <cstring>
#pragma pack(push, 1)
struct PcapGlobalHeader {
uint32_t magic_number;
uint16_t version_major;
uint16_t version_minor;
int32_t thiszone;
uint32_t sigfigs;
uint32_t snaplen;
uint32_t network;
};
#pragma pack(pop)
uint16_t swap16(uint16_t value) {
return (value >> 8) | (value << 8);
}
uint32_t swap32(uint32_t value) {
return ((value & 0x000000FFU) << 24) |
((value & 0x0000FF00U) << 8) |
((value & 0x00FF0000U) >> 8) |
((value & 0xFF000000U) >> 24);
}
二、循环读取记录头,避免整包读入带来的内存压力
全局头之后,文件会重复出现多个数据包记录。每个记录头固定为 16 字节,分别是抓包时间秒、微秒、实际保存长度和原始包长度。需要特别注意的是 incl_len 和 orig_len 可能不同,很多抓包工具为了控制文件体积会按 snaplen 截断报文,因此有效数据必须是 incl_len,而不是 orig_len。
从文件 IO 角度看,解析器应当以流式方式循环读取,而不是一次性把整个 pcap 文件塞进内存。大流量分析场景下 pcap 文件动辄几十 GB,逐块读取能够保持较低的内存占用。每次读取记录头后,根据 incl_len 分配一块缓冲区,读入对应长度的原始报文,再交给后续协议解析函数。
代码示例中加入了 EOF 和读取长度校验。遇到损坏的截断文件时,gcount 返回值可能与预期不相等,这时应当及时退出,而不是继续使用不完整的数据。时间戳可以转换为 timeval 或人类可读时间,方便后续按时间窗口做聚合。
#include <fstream>
#include <vector>
#include <cstdint>
#pragma pack(push, 1)
struct PcapRecordHeader {
uint32_t ts_sec;
uint32_t ts_usec;
uint32_t incl_len;
uint32_t orig_len;
};
#pragma pack(pop)
bool parsePcapFile(const char* path) {
std::ifstream input(path, std::ios::binary);
if (!input.is_open()) {
return false;
}
PcapGlobalHeader global_header;
input.read(reinterpret_cast<char*>(&global_header), sizeof(global_header));
if (!input) {
return false;
}
bool need_swap = false;
if (global_header.magic_number == 0xA1B2C3D4) {
need_swap = false;
} else if (global_header.magic_number == 0xD4C3B2A1) {
need_swap = true;
} else {
return false;
}
while (input.peek() != EOF) {
PcapRecordHeader record;
input.read(reinterpret_cast<char*>(&record), sizeof(record));
if (input.gcount() != sizeof(record)) {
break;
}
if (need_swap) {
record.ts_sec = swap32(record.ts_sec);
record.ts_usec = swap32(record.ts_usec);
record.incl_len = swap32(record.incl_len);
record.orig_len = swap32(record.orig_len);
}
std::vector<unsigned char> packet(record.incl_len);
input.read(reinterpret_cast<char*>(packet.data()), packet.size());
if (!input) {
break;
}
// 在此处调用 ethernet、ip、tcp 解析函数
}
return true;
}
三、libpcap 的离线解析接口与手动实现的差异
对很多生产环境来说,直接使用 libpcap 或 Npcap 的离线读取接口是更省事的方案。pcap_open_offline 能自动处理字节序、读取被截断的报文、识别不同链路层类型,并且提供 BPF 过滤表达式支持。在需要快速过滤特定 IP 或端口时,用 pcap_compile 和 pcap_setfilter 比手动判断效率高得多。
不过 libpcap 也有边界。它封装掉了底层文件布局细节,当排查文件损坏、分析不同版本 pcapng 差异,或者需要在没有该库的嵌入式环境里解析时,手动方式仍然不可替代。实际项目中可以将两种方式结合:手动解析用于格式校验和调试,libpcap 用于主业务流程。
下面这段代码演示了使用 libpcap 读取离线 pcap 文件的基本流程。pcap_next_ex 返回 1 表示成功读到包,0 表示超时,负数表示错误或 EOF。每个包的 caplen 和 len 分别对应记录头里的 incl_len 和 orig_len。
#include <pcap.h>
#include <cstdio>
void readWithLibpcap(const char* path) {
char errbuf[PCAP_ERRBUF_SIZE];
pcap_t* handle = pcap_open_offline(path, errbuf);
if (handle == nullptr) {
std::fprintf(stderr, "open pcap failed: %s\n", errbuf);
return;
}
struct pcap_pkthdr* header = nullptr;
const u_char* packet = nullptr;
int ret = 0;
while ((ret = pcap_next_ex(handle, &header, &packet)) >= 0) {
if (ret == 0) {
continue;
}
std::printf("caplen=%u len=%u ts=%ld.%06ld\n",
header->caplen,
header->len,
header->ts.tv_sec,
header->ts.tv_usec);
}
pcap_close(handle);
}
四、从以太网到 TCP 负载:计算偏移与校验长度
拿到单包的原始数据后,解析方向通常从链路层开始。以常见的以太网为例,前 14 字节是目的 MAC、源 MAC 和 EtherType。EtherType 为 0x0800 时后续是 IPv4 报文,0x86DD 则是 IPv6。如果存在 VLAN 标签,EtherType 位置会变成 0x8100,需要额外跳过 4 字节再做判断。
IPv4 头部最短 20 字节,首字节的低 4 位表示 IHL,单位是 4 字节。通过 IHL 可以计算出实际 IP 头长度,进而定位到 TCP 或 UDP 头。网络序字段在做比较和输出时,必须使用 ntohs 或 ntohl 转换。下面示例只截取 TCP 五元组中的源 IP、目的 IP、源端口和目的端口,实际业务中还会继续读取 TCP 序列号、标志位和负载偏移。
解析代码里每个长度判断都不能省略。抓包文件里的包可能被截断,也可能来自非以太网链路。直接按固定偏移访问数组而不检查大小,是 pcap 解析器崩溃和堆破坏的最常见原因。通过层层校验长度,可以在数据异常时直接丢弃,避免让一个畸形包影响整个批量分析任务。
#include <cstdint>
#include <cstring>
#include <vector>
#include <arpa/inet.h> // Linux / macOS 下使用
// Windows 下可改用 winsock2.h,并链接 ws2_32.lib
struct EthernetHeader {
uint8_t dst_mac[6];
uint8_t src_mac[6];
uint16_t ether_type;
};
struct Ipv4Header {
uint8_t version_ihl;
uint8_t tos;
uint16_t total_length;
uint16_t identification;
uint16_t flags_fragment;
uint8_t ttl;
uint8_t protocol;
uint16_t header_checksum;
uint32_t src_addr;
uint32_t dst_addr;
};
void parseTcpPacket(const std::vector<unsigned char>& data) {
if (data.size() < sizeof(EthernetHeader)) {
return;
}
EthernetHeader eth;
std::memcpy(ð, data.data(), sizeof(eth));
uint16_t ether_type = ntohs(eth.ether_type);
if (ether_type != 0x0800) {
return;
}
if (data.size() < sizeof(EthernetHeader) + sizeof(Ipv4Header)) {
return;
}
Ipv4Header ipv4;
std::memcpy(&ipv4, data.data() + sizeof(EthernetHeader), sizeof(ipv4));
uint8_t ihl = (ipv4.version_ihl & 0x0F) * 4;
uint8_t protocol = ipv4.protocol;
if (protocol != 6) { // 6 表示 TCP
return;
}
size_t tcp_offset = sizeof(EthernetHeader) + ihl;
if (data.size() < tcp_offset + 20) {
return;
}
uint16_t src_port = (data[tcp_offset] << 8) | data[tcp_offset + 1];
uint16_t dst_port = (data[tcp_offset + 2] << 8) | data[tcp_offset + 3];
uint32_t src_ip = ipv4.src_addr;
uint32_t dst_ip = ipv4.dst_addr;
// src_port、dst_port、src_ip、dst_ip 就是五元组中的四元信息
}
五、解析中容易忽视的若干细节
除了全局头字节序,结构体对齐是另一个隐蔽问题。如果使用 pragma pack 定义结构体,可以保证在主流编译器上按紧凑布局读取,但跨编译器或平台时仍然建议依靠 memcpy 和手工偏移。这样代码更容易移植到 ARM 等对非对齐访问敏感的架构上。
时间戳的处理也要结合业务需求。ts_sec 和 ts_usec 是相对 UTC 的秒和微秒,如果要写入数据库,可以转换成 64 位微秒时间戳;如果要在前端展示,则可以用 gmtime_r 或 localtime_r 转换。抓包工具写入的 thiszone 字段一般不可靠,不建议把它作为时区依据。
最后还应当区分 pcap 与 pcapng。本文描述的是传统 pcap 格式,文件尾不包含索引。而 Wireshark 新版默认会生成 pcapng,后者支持多接口、包注释和自定义块,结构更复杂。如果遇到无法识别的 magic number,大概率是 pcapng,需要引入 pcapng 解析库或使用 libpcap 1.10 以上的能力。