EADST

Extract Webpage Information with Python

Here is the python program to extract webpage information with BeautifulSoup and save the data in a CSV file.

from bs4 import BeautifulSoup
import urllib.request
import pandas as pd

url = 'file:///Users/xd/Desktop/ieee/Region_5_Student_Branch_Counselors_and_Chairs.htm'
save_file = 'ieee_info_1'
html = urllib.request.urlopen(url).read()

soup = BeautifulSoup(html, "html.parser")

universities = soup.find_all('div', class_='spoName bullet pad-t15')
people = soup.find_all('div', class_='roster-results')

for u, p in zip(universities, people):
    info = p.find_all('p')
    university = u.get_text()
    name = info[0].get_text()
    if name == 'Position Vacant':
        continue
    title = info[2].get_text()
    address = info[3].get_text() + ', ' + info[4].get_text()
    email = info[-1].get_text()[7:]

    content = [[university, name, title, address, email]]
    list_name = ['university', 'name', 'title', 'address', 'email']
    data = pd.DataFrame(columns=list_name, data=content)
    data.to_csv("{}.csv".format(save_file), mode='a', index=False, header=False, encoding='utf-8')
About Me
XD
Goals determine what you are going to be.
Category
标签云
GPTQ 搞笑 LLAMA RL torchinfo Web Firewall 净利润 Streamlit PyCharm UNIX Card TensorRT 论文 CTC Safetensors XML ChatGPT 签证 Claude Github Password FP32 ModelScope diffusers Website Vim TTS Image2Text HuggingFace logger Bin Numpy Qwen 财报 Heatmap CLAP DeepSeek TensorFlow Review Llama Diagram XGBoost Git GPT4 WAN Excel FP16 Sklearn Plotly 阿里云 llama.cpp Jupyter 论文速读 MD5 图形思考法 NLP tar Qwen2 Datetime Input PDB Distillation GIT FlashAttention Bipartite QWEN transformers PIP Qwen2.5 Video Plate 顶会 关于博主 Statistics DeepStream Harness Pytorch git tqdm Miniforge CSV GGML Jev Algorithm Shortcut Use 腾讯云 Magnet JSON Hungarian Baidu ms-swift OpenAI Animate Agent NameSilo RGB Pandas VPN CC Hotel PDF 证件照 Mixtral FP8 FastAPI Docker Hilton VSCode IndexTTS2 Freesound 云服务器 uwsgi Tracking Math NLTK Translation Google v0.dev ONNX Pickle hf PyTorch 第一性原理 API网关 CEIR YOLO 图标 Interview Random Template Color Michelin Linux Permission 音频 Transformers UI News CUDA SQL Proxy Jetson OpenCV LLM Rebuttal Augmentation OCR 继承 Quantization icon 多线程 Windows v2ray Crawler Search HaggingFace Anaconda Clash BF16 API 飞书 VGG-16 GoogLeNet Disk 域名 强化学习 算法题 CV LoRA Attention Data LaTeX Paper Quantize InvalidArgumentError Tiktoken Python Tensor Django Conda Bert Logo Ubuntu 报税 Bitcoin COCO SPIE 递归学习法 SAM Land Base64 scipy C++ printf CAM Breakpoint BeautifulSoup mmap Food 版权 多进程 Gemma Pillow SVR LeetCode Ptyhon EXCEL SQLite FP64 Zip Paddle RAR Markdown TSV uWSGI Cloudreve Vmess AI ResNet-50 Dataset WebCrawler 公式 Domain git-lfs BTC Knowledge Nginx
站点统计

本站现有博文337篇,共被浏览961562次

本站已经建立2673天!

热门文章
文章归档
回到顶部