爬虫（5）爬取多页数据 - 代码天地

爬虫（5）爬取多页数据

其他 2021-03-24 03:15:45 阅读次数: 0

我们点开其他年份的GDP数据时，会发现网站的变化只有后面的数字变成了相应的年份，所以我们可以通过for循环来实现对多页数据的爬取

from selenium import webdriver
from bs4 import BeautifulSoup
import csv

driver=webdriver.Chrome()
out=open('d:/gdp.csv','w',newline='')
csv_write=csv.writer(out,dialect='excel')
for year in range(1960,2020):
    url="https://www.kylc.com/stats/global/yearly/g_gdp/%d.html"%year
    xpath="/html/body/div[2]/div[1]/div[5]/div[1]/div/div/div/table"
    driver.get(url)
    tablel=driver.find_element_by_xpath(xpath).get_attribute('innerHTML')

    soup=BeautifulSoup(tablel,"html.parser")
    table=soup.find_all('tr')
    for row in table:
        cols=[col.text for col in row.find_all('td')]
        if len(cols)==0 or not cols[0].isdigit():
            continue
        cols.append(year)
        csv_write.writerow(cols)
out.close()
driver.close()

这里要先注意，打开文件只运行一遍，也就是要放到for循环外面，不然新数据会覆盖掉原数据
年份是从1960年到2019年，所以range参数到2020
同时url也需要更改，改为

url="https://www.kylc.com/stats/global/yearly/g_gdp/%d.html"%year

为了不混淆年份对应的GDP，可以在最后面加上一个年份

自此，爬取GDP数据结束

猜你喜欢

转载自blog.csdn.net/qq_53029299/article/details/114851722

爬虫（5）爬取多页数据

Python 爬虫爬取多页数据

Python爬虫——爬取网站多页数据

爬虫——爬取网页数据存入表格

利用爬虫爬取简单页码类网页数据

爬虫03_基于requests的分页数据的爬取

Scrapy 实现爬取多页数据 + 多层url数据爬取

Python爬虫项目：爬虫爬取BeautifulSoup模块分析网页数据

爬虫（5）：爬取拉钩网数据

JAVA爬虫爬取网页数据数据库中,并且去除重复数据

正则爬取网页数据(二)

正则爬取网页数据(三)

Python爬取网页数据

爬取雪球房产数据随意页数

爬取网页数据python

java网页数据爬取

python初学-爬取网页数据

如何快速爬取网页数据

Scrapy爬取网页数据

jsoup爬取网页数据

使用 Python 爬取网页数据

Java爬取网页数据

爬取网页数据基础

python爬取网页数据方法

使用XPath爬取网页数据

node爬取cnode首页数据

Python 简单爬取网页数据

Python3.5-爬虫实战-爬取网页数据并且导入excel

你以为Python爬虫只能爬取网页数据吗？APP也是可以的呢！

接着上次的python爬虫，今天进阶一哈，局部解析爬取网页数据

今日推荐

周排行

LRU cache算法

windows10, 自带的OpenSSH, key权限问题, 文件权限问题

测试用例书写方法

HIVE-默认分隔符的（linux系统的特殊字符）查看，输入和修改

最贵的AMD 7nm显卡来了！这设计够狂野

java多线程简单demo

[ 转载 ]在Android系统上使用busybox——最简单的方法

QT connect学习

BFSIFT算法分析

Xcode10：library not found for -lstdc++.6.0.9 临时解决

每日归档

更多

2024-08-06(0)

2024-08-05(0)

2024-08-04(0)

2024-08-03(0)

2024-08-02(0)

2024-08-01(0)

2024-07-31(0)

2024-07-30(0)

2024-07-29(0)

2024-07-28(0)