Python进阶教程:网络编程与数据抓取 目录Python进阶教程网络编程与数据抓取一、HTTP 基础二、urllib标准库 HTTP 客户端三、requests更优雅的 HTTP 库四、HTML 解析BeautifulSoup五、实战抓取网页文章标题六、爬虫的注意事项七、进阶接口调用API总结Python进阶教程网络编程与数据抓取本文是Python 入门教程系列的第 5 篇。前面四篇介绍了基础语法、OOP、文件操作、常用标准库本篇介绍网络编程与数据抓取爬虫基础。一、HTTP 基础网络编程的核心是 HTTP 协议。HTTP 请求主要由四部分组成方法GET获取、POST提交、PUT、DELETE 等URL资源地址请求头User-Agent、Cookie、Content-Type 等请求体POST 时携带的数据响应同样包含状态码200 成功、404 不存在、500 服务器错误、响应头和响应体。二、urllib标准库 HTTP 客户端importurllib.requestimporturllib.parse# GET 请求urlhttps://httpbin.org/getrequrllib.request.Request(url,headers{User-Agent:Mozilla/5.0})withurllib.request.urlopen(req,timeout10)asresp:print(resp.status)# 200print(resp.read().decode(utf-8)[:200])# POST 请求dataurllib.parse.urlencode({name:Alice,age:20}).encode()requrllib.request.Request(https://httpbin.org/post,datadata)withurllib.request.urlopen(req)asresp:print(resp.read().decode(utf-8)[:200])三、requests更优雅的 HTTP 库requests 是第三方库pip install requests是实际开发中的首选importrequests# GET 请求resprequests.get(https://httpbin.org/get,params{q:python},timeout10)print(resp.status_code)# 200print(resp.json())# 自动解析 JSON# POST 请求resprequests.post(https://httpbin.org/post,json{name:Alice})print(resp.json())# 自定义请求头模拟浏览器headers{User-Agent:Mozilla/5.0 (Windows NT 10.0; Win64; x64)}resprequests.get(https://httpbin.org/headers,headersheaders)# 下载文件resprequests.get(https://httpbin.org/image/png,streamTrue)withopen(image.png,wb)asf:forchunkinresp.iter_content(chunk_size8192):f.write(chunk)四、HTML 解析BeautifulSoup抓取网页后需要解析 HTMLBeautifulSoup 是最常用的工具pip install beautifulsoup4frombs4importBeautifulSoupimportrequests html htmlbody h1Python 教程/h1 div classarticle a href/p1第一篇/a a href/p2第二篇/a /div /body/html soupBeautifulSoup(html,html.parser)# 获取标题print(soup.h1.text)# Python 教程# 按 class 查找divsoup.find(div,class_article)# 查找所有链接foraindiv.find_all(a):print(a.text,a[href])# 第一篇 /p1# 第二篇 /p2五、实战抓取网页文章标题综合运用以上知识写一个抓取网页所有链接和标题的小工具importrequestsfrombs4importBeautifulSoupdeffetch_links(url):抓取页面中所有链接及其文本try:headers{User-Agent:Mozilla/5.0}resprequests.get(url,headersheaders,timeout10)resp.raise_for_status()# 非 200 会抛出异常soupBeautifulSoup(resp.text,html.parser)links[]forainsoup.find_all(a,hrefTrue):texta.text.strip()or(无文本)links.append((text[:30],a[href]))returnlinksexceptrequests.RequestExceptionase:print(f请求失败{e})return[]# 使用示例urlhttps://example.comfortext,hrefinfetch_links(url)[:10]:print(f{text}-{href})六、爬虫的注意事项合法、规范的爬虫需要注意遵守 robots.txt访问站点前检查https://站点/robots.txt了解允许爬取的内容。控制请求频率用 time.sleep 间隔请求避免给服务器造成压力。设置合理 UA识别为真实浏览器但不要伪装成他人。尊重版权只抓取允许的数据注意使用条款。反爬处理遇到验证码、登录墙时不要强行绕过。importtimeimportrequests urls[https://httpbin.org/get]*5forurlinurls:resprequests.get(url,timeout10)print(resp.status_code)time.sleep(2)# 每 2 秒请求一次礼貌抓取七、进阶接口调用API现代开发更多是调用 API 获取 JSON 数据配合上篇的 json 库非常方便importrequestsimportjsondefcall_api(url,paramsNone):resprequests.get(url,paramsparams,timeout10)ifresp.status_code200:returnresp.json()else:print(fAPI 返回错误{resp.status_code})returnNone# 调用公开 API 获取天气信息示例datacall_api(https://httpbin.org/json)ifdata:print(json.dumps(data,ensure_asciiFalse,indent2))总结本篇介绍了 HTTP 基础、urllib 与 requests 两种 HTTP 客户端、BeautifulSoup 网页解析、以及爬虫的规范与注意事项并提供了两个实战工具。下一篇将介绍多线程与多进程敬请期待

相关新闻

最新新闻

SerenityOS 命令行选项解析指南:getopt 与 getopt_long 用法、返回值与底层实现

SerenityOS 命令行选项解析指南:getopt 与 getopt_long 用法、返回值与底层实现

SerenityOS 命令行选项解析指南:getopt 与 getopt_long 用法、返回值与底层实现 【免费下载链接】serenity The Serenity Operating System 🐞 项目地址: https://gitcode.com/GitHub_Trending/se/serenity 导读 本文以 getopt(3) 手册 为核心&a…

2026/10/5 3:18:56
轻量服务器还是ECS?大促云服务器选购与避坑实战指南

轻量服务器还是ECS?大促云服务器选购与避坑实战指南

每年大促节点,群里永远有人在问同一个问题:“38元的轻量服务器到底怎么抢?为什么我每次点进去都是已售罄?68元直购和99元的ECS我到底选哪个?”作为一个常年帮团队和自己采购云服务器的老用户,我太清楚这种纠…

2026/10/5 3:42:18
为 AI 代理的 Review 动作编写 Cedar 审批门控策略:review-agent-governance 策略编写实战指南

为 AI 代理的 Review 动作编写 Cedar 审批门控策略:review-agent-governance 策略编写实战指南

为 AI 代理的 Review 动作编写 Cedar 审批门控策略:review-agent-governance 策略编写实战指南 【免费下载链接】agents Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, and Google Antigravity 项目地址:…

2026/10/5 19:39:38
PaddleOCR 手写数学公式识别算法 CAN 实战指南:Counting-Aware Network 训练、评估与推理部署

PaddleOCR 手写数学公式识别算法 CAN 实战指南:Counting-Aware Network 训练、评估与推理部署

PaddleOCR 手写数学公式识别算法 CAN 实战指南:Counting-Aware Network 训练、评估与推理部署 【免费下载链接】PaddleOCR Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between i…

2026/10/5 16:06:34
Spring源码解析:构造器注入的类型转换与候选匹配机制

Spring源码解析:构造器注入的类型转换与候选匹配机制

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/10/5 5:51:09
openai-agents-python 多模型接入指南:深入解析 AnyLLMModel 适配层与 any-llm 路由

openai-agents-python 多模型接入指南:深入解析 AnyLLMModel 适配层与 any-llm 路由

openai-agents-python 多模型接入指南:深入解析 AnyLLMModel 适配层与 any-llm 路由 【免费下载链接】openai-agents-python A lightweight, powerful framework for multi-agent workflows 项目地址: https://gitcode.com/GitHub_Trending/op/openai-agents-pyth…

2026/10/5 5:40:36

日新闻

周新闻

月新闻