EADST

Python: Obtain Baidu Images Using Web Crawler

Python: Obtain Baidu Images Using Web Crawler.

Here is the main code.

# -- coding: utf-8 --
import os
import re
import time
import requests

class CarCollect():

def __init__(self, path='./name.txt'):
    self.num = 1
    self.class_number = 0
    self.line_list = []
    with open(path, encoding='utf-8') as file:
        self.line_list = [k.strip() for k in file.readlines()]
        self.class_number = int(self.line_list[0])
        self.line_list = self.line_list[1:]

def dowmload_picture(self, html, keyword, save_path):
    pic_url = re.findall('"objURL":"(.*?)",', html, re.S)  # get image url
    print('Finding keyword: ' + keyword + ' images, start downloading...')
    for each in pic_url:
        print('='*60)
        print('Downloading ' + keyword + ' number ' + str(self.num) + ' image, url: ' + str(each))
        try:
            if each:
                pic = requests.get(each, timeout=7)
                string = save_path + r'\\' + keyword + '_' + str(self.num) + '.jpg'
                if len(pic.content) > 10000: # img size > 10k
                    with open(string, 'wb') as fp:
                        fp.write(pic.content)
                        self.num += 1
        except BaseException:
            print('error, cannot download')
        if self.num > self.class_number:
            break

def __call__(self):
    headers = {
        'Accept-Language': 'zh-CN,zh;q=0.8,zh-TW;q=0.7,zh-HK;q=0.5,en-US;q=0.3,en;q=0.2',
        'Connection': 'keep-alive',
        'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64; rv:60.0) Gecko/20100101 Firefox/60.0',
        'Upgrade-Insecure-Requests': '1'
    }
    session = requests.Session()
    session.headers = headers

    for word in self.line_list:
        # create a folder
        save_path = word + '_file'
        time_now = time.strftime("%Y%m%d_%H%M%S", time.localtime())
        save_path += "_" + time_now
        os.mkdir(save_path)
        # get images
        url = 'https://image.baidu.com/search/flip?tn=baiduimage&ie=utf-8&word=' + word + '&pn='
        image_number = 0
        self.num = 1
        while image_number < self.class_number:
            try:
                result = session.get(url + str(image_number), timeout=10, allow_redirects=False)
                self.dowmload_picture(result.text, word, save_path)
            except:
                print('Internet error')
            image_number += 60

if name == 'main': path = './keywords.txt' car_collect = CarCollect(path) car_collect() print('Done.')

Here is the text file, keywords.txt. The first line is the number we want to obtain from each keyword. The following lines are the keywords.

20
Dog
Cat

相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
Qwen2 BF16 腾讯云 Sklearn Jupyter Django SPIE WebCrawler MD5 PyCharm Clash OpenAI Bipartite Tiktoken TSV Conda uWSGI Proxy GGML PDF 关于博主 强化学习 Website diffusers Card Baidu CC GPT4 Diagram NLTK Google Zip mmap Heatmap Qwen PDB Quantization C++ uwsgi Dataset Shortcut 图形思考法 torchinfo transformers RGB Git HaggingFace Color FP32 证件照 Password TensorFlow Breakpoint CSV XML COCO Input UNIX RAR Pytorch OCR Review Anaconda FP16 Magnet 版权 Plate 音频 XGBoost Bert 财报 Datetime scipy Tensor GIT LLAMA llama.cpp FP8 Distillation LLM Numpy Cloudreve 算法题 printf Vim ms-swift PyTorch ResNet-50 Translation Image2Text ChatGPT git-lfs WAN Use Safetensors Random NLP logger Bitcoin 飞书 Knowledge BTC Freesound Pillow FlashAttention Vmess Video ONNX Pandas SQLite Augmentation Llama 净利润 icon v2ray Plotly 论文 Miniforge git SVR Disk Paper Paddle CLAP 阿里云 Mixtral LoRA Pickle Python RL 图标 Jetson News Linux 签证 公式 Attention GoogLeNet Search CTC AI 多线程 Hotel 云服务器 v0.dev tar Nginx VPN Hilton Claude 报税 Ubuntu Qwen2.5 Algorithm VSCode Michelin DeepSeek TTS Excel QWEN Interview ModelScope Base64 LeetCode Bin Statistics Github 论文速读 GPTQ LaTeX Agent JSON 多进程 hf 继承 SQL 第一性原理 InvalidArgumentError UI Hungarian Logo Data Markdown Tracking YOLO IndexTTS2 顶会 Crawler Math 搞笑 DeepStream Food Transformers FastAPI Streamlit Gemma Permission EXCEL HuggingFace NameSilo tqdm Web API Domain CAM Docker PIP Template Firewall 域名 FP64 CEIR Land SAM CUDA CV Windows Rebuttal 递归学习法 VGG-16 Ptyhon Quantize BeautifulSoup TensorRT Animate OpenCV
站点统计

本站现有博文333篇,共被浏览922615

本站已经建立2628天!

热门文章
文章归档
回到顶部