Python B 文件里的每行数据匹配 A 文件里的内容 逻辑问题
A 文件是个 json 文件: a.txt
{ "_id": "113.254.82.124", "_index": "fofapro_subdomain", "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } { "_id": "http://10.254.82.12", "_index": "fofapro_subdomain", "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } { "_id": "https://192.168.1.10:9090", "_index": "fofapro_subdomain", "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } { "_id": "127.0.0.1:8343", "_index": "fofapro_subdomain", "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", }
B 文件: b.txt
127.0.01 192.168.1.10 192.168.88.88
代码
import re import json def filesJson(filepath,dstpaths): datas = set() #正则匹配 rule = re.compile('^[a-zA-z]{1}.*$') with open(filepath, 'r', encoding='UTF-8') as a, open(dstpaths, 'r', encoding='UTF-8') as b: b.seek(0) for realine_a in a: json_datas = json.loads(realine_a) ips = json_datas['_id'] if rule.findall(ips): ips = ips.strip("http[s]?://") ips = ips.split(":")[0] datas.add(ips) for realine_b in b: if realine_b in datas: print(realine_b) else: break if __name__ == '__main__': file_paths = "a.txt" dstpaths = 'b.txt' filesJson(file_paths, dstpaths)
我的想法是把 A 文件里的 IP,去除协议和端口,只保留 IP 写入到一个集合中,然后在通过 B 文件的数据去匹配这个集合,有没有这个 IP,如果有这个 IP,把 A 文件这行数据写入到 C 文件中,现在问题是 B 文件无法匹配 A 文件的数据,而且如果 A 文件内容是几百万行数据,B 文件内容是几万行数据,这种逻辑是不是有很大的问题。