跳至主要內容
  • Hostloc 空間訪問刷分
  • 售賣場
  • 廣告位
  • 賣站?

4563博客

全新的繁體中文 WordPress 網站
  • 首頁
  • Python B 文件里的每行数据匹配 A 文件里的内容 逻辑问题
未分類
28 5 月 2020

Python B 文件里的每行数据匹配 A 文件里的内容 逻辑问题

Python B 文件里的每行数据匹配 A 文件里的内容 逻辑问题

資深大佬 : aaa5838769 1

A 文件是个 json 文件: a.txt

{  "_id": "113.254.82.124",  "_index": "fofapro_subdomain",  "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } {  "_id": "http://10.254.82.12",  "_index": "fofapro_subdomain",  "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } {  "_id": "https://192.168.1.10:9090",  "_index": "fofapro_subdomain",  "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } {  "_id": "127.0.0.1:8343",  "_index": "fofapro_subdomain",  "header": "HTTP/1.1 401 UnauthorizedrnConnection: closernContent-Length: 195rnCache-Control: no-cachernContent-Type: text/htmlrnDate: Sat, 20 Oct 2018 15:59:44 GMTrnEtag: "0-29d-b90"rnServer: Embedthis-Appweb/3.3.1rnWww-Authenticate: Basic realm="DCS-2530L"rnX-Frame-Options: SAMEORIGINrn", } 

B 文件: b.txt

127.0.01 192.168.1.10 192.168.88.88 

代码

import re import json  def filesJson(filepath,dstpaths):     datas = set()     #正则匹配     rule = re.compile('^[a-zA-z]{1}.*$')     with open(filepath, 'r', encoding='UTF-8') as a, open(dstpaths, 'r', encoding='UTF-8') as b:         b.seek(0)         for realine_a in a:             json_datas = json.loads(realine_a)             ips = json_datas['_id']             if rule.findall(ips):                 ips = ips.strip("http[s]?://")                 ips = ips.split(":")[0]                 datas.add(ips)         for realine_b in b:             if realine_b in datas:                 print(realine_b)             else:                 break  if __name__ == '__main__':     file_paths = "a.txt"     dstpaths = 'b.txt'     filesJson(file_paths, dstpaths) 

我的想法是把 A 文件里的 IP,去除协议和端口,只保留 IP 写入到一个集合中,然后在通过 B 文件的数据去匹配这个集合,有没有这个 IP,如果有这个 IP,把 A 文件这行数据写入到 C 文件中,现在问题是 B 文件无法匹配 A 文件的数据,而且如果 A 文件内容是几百万行数据,B 文件内容是几万行数据,这种逻辑是不是有很大的问题。

大佬有話說 (20)

  • 資深大佬 : liprais

    pyspark 安排了

  • 主 資深大佬 : aaa5838769

    @liprais 感谢回复,但是我现在的代码逻辑是无法匹配到数据。

  • 資深大佬 : telnetning

    确认一下每个 Json 是在一行吗; rule 的正则可以再看一看?;可以直接断点调试一下~

  • 資深大佬 : noqwerty

    别的不说,你现在每插入一条数据就要遍历一次 b 完全没意义啊,拿到两个 set 之后直接取交集不就行了吗

  • 主 資深大佬 : aaa5838769

    @noqwerty 我试过交集,发现还是没有数据。

  • 主 資深大佬 : aaa5838769

    @telnetning 每个 json 不在一行,一个 json 有很多数据,我只是提取出来几个字段,我的正则主要是为了匹配到 http 开头的,然后删除协议,如果有端口,进行截取,再存到集合中。

  • 資深大佬 : AFlash

    试试把“else:break”去掉,

  • 資深大佬 : noqwerty

    @aaa5838769 #6 不在一行的话你的 json.loads(realine_a) 不会报错吗?

  • 資深大佬 : ipwx

    @aaa5838769 。。。

    b_set = set(filter(lambda s: s, (l.strip() for l in open(‘b.txt’, ‘r’, encoding=’utf-8′))))
    with open(‘a.txt’, ‘r’, encoding=’utf-8′) as a_file, open(‘c.txt’, ‘r’, encoding=’utf-8′) as c_file:
    ….for s in a:
    ……..do what you have do to get s_ip
    ……..if s_ip in b_set:
    …………c_file.write(s_ip + ‘n’)

  • 資深大佬 : ipwx

    open(‘c.txt’, ‘r’, encoding=’utf-8′) as c_file => open(‘c.txt’, ‘w’, encoding=’utf-8′) as c_file:

  • 資深大佬 : ipwx

    主你这需求,像我上面这样,把 b 做成 set 过一遍 a 就行了。不理解你在纠结啥。

  • 資深大佬 : xujunfu

    取出 b 文件所以 ips,再把每个 ip 转成 long 数值,将数组排序;遍历 a 文件取出 ip 并转 long 数值,二分查找在不在 long_ips 数组中

  • 主 資深大佬 : aaa5838769

    @noqwerty 不好意思,在一行上,我脑子短路啊,我专门格式化一下,然后复制到论坛里。

  • 主 資深大佬 : aaa5838769

    @ipwx 我之前用过两个 set 做交集,也没有匹配到数据,所以换了另一种逻辑。

  • 資深大佬 : wuwukai007

    a 文件格式是这样?
    {‘_id’:xxx}下一行
    {‘_id’:xxx}下一行

  • 主 資深大佬 : aaa5838769

    @wuwukai007 是的

  • 資深大佬 : rrfeng

    这个最佳实践是:

    b 预读完,做成 hash ( dict ),ip 当 key,放在内存里。

    然后 a 用流式 json 处理解析完直接查找 hash 表。

  • 資深大佬 : wuwukai007

    a.txt 格式
    {xxxxxxxx}n
    {xxxxxxxx}n
    b.txt 格式
    192.168.1.10
    192.168.1.10

    import pandas as pd
    import re
    df = pd.read_csv(r’a.txt’,header=None)
    df.rename(columns={0:’ip’},inplace=True)
    regex = ‘(([01]{0,1}d{0,1}d|2[0-4]d|25[0-5]).){3}([01]{0,1}d{0,1}d|2[0-4]d|25[0-5])’
    df[‘ip_host’] = df.ip.str.replace(rf’^((?!{regex}).)*’,”).str.replace(‘(:d+)|”‘,”)
    df2 = pd.read_csv(r”b.txt”,sep=’n’,header=None)
    df2.columns = [‘ip_host’]
    df3 = df2.merge(df,on=’ip_host’,how=’inner’)
    df3.drop(‘ip_host’,axis=1,inplace=True)
    df3.to_csv(r”c.txt”,index=None,header=None,quoting=None,quotechar=”‘”)

  • 資深大佬 : wuwukai007

    a.txt 格式
    {xxxxxxxx}换行符
    {xxxxxxxx}换行符
    b.txt 格式
    192.168.1.10 换行符
    192.168.1.10 换行符

    import pandas as pd
    import re
    df = pd.read_csv(r’a.txt’,header=None)
    df.rename(columns={0:’ip’},inplace=True)
    regex = ‘(([01]{0,1}d{0,1}d|2[0-4]d|25[0-5]).){3}([01]{0,1}d{0,1}d|2[0-4]d|25[0-5])’
    df[‘ip_host’] = df.ip.str.replace(rf’^((?!{regex}).)*’,”).str.replace(‘(:d+)|”‘,”)
    df2 = pd.read_csv(r”b.txt”,sep=’n’,header=None)
    df2.columns = [‘ip_host’]
    df3 = df2.merge(df,on=’ip_host’,how=’inner’)
    df3.drop(‘ip_host’,axis=1,inplace=True)
    df3.to_csv(r”c.txt”,index=None,header=None,quoting=None,quotechar=”‘”)

    输出 c.txt
    {xxxxxxxx}换行符
    {xxxxxxxx}换行符

  • 主 資深大佬 : aaa5838769

    @rrfeng value 为空么?

文章導覽

上一篇文章
下一篇文章

AD

其他操作

  • 登入
  • 訂閱網站內容的資訊提供
  • 訂閱留言的資訊提供
  • WordPress.org 台灣繁體中文

51la

4563博客

全新的繁體中文 WordPress 網站
返回頂端
本站採用 WordPress 建置 | 佈景主題採用 GretaThemes 所設計的 Memory
4563博客
  • Hostloc 空間訪問刷分
  • 售賣場
  • 廣告位
  • 賣站?
在這裡新增小工具