200字范文,内容丰富有趣,生活中的好帮手!
200字范文 > python爬虫获取网易云音乐歌单

python爬虫获取网易云音乐歌单

时间:2019-02-28 01:40:22

相关推荐

python爬虫获取网易云音乐歌单

代码如下:

from bs4 import BeautifulSoupimport requestsimport timeheaders = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36'}for i in range(0, 1330, 35):print(i)time.sleep(2)url = '/discover/playlist/?cat=欧美&order=hot&limit=35&offset=' + str(i)response = requests.get(url=url, headers=headers)html = response.textsoup = BeautifulSoup(html, 'html.parser')# 获取包含歌单详情页网址的标签ids = soup.select('.dec a')# 获取包含歌单索引页信息的标签lis = soup.select('#m-pl-container li')print(len(lis))for j in range(len(lis)):# 获取歌单详情页地址url = ids[j]['href']# 获取歌单标题title = ids[j]['title']# 获取歌单播放量play = lis[j].select('.nb')[0].get_text()# 获取歌单贡献者名字user = lis[j].select('p')[1].select('a')[0].get_text()# 输出歌单索引页信息print(url, title, play, user)# 将信息写入CSV文件中with open('playlist.csv', 'a+', encoding='utf-8-sig') as f:f.write(url + ',' + title + ',' + play + ',' + user + '\n')

获取歌单信息如下:

二次代码如下:

from bs4 import BeautifulSoupimport pandas as pdimport requestsimport timedf = pd.read_csv('playlist.csv', header=None, error_bad_lines=False, names=['url', 'title', 'play', 'user'])headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36'}for i in df['url']:time.sleep(2)url = '' + iresponse = requests.get(url=url, headers=headers)html = response.textsoup = BeautifulSoup(html, 'html.parser')# 获取歌单标题title = soup.select('h2')[0].get_text().replace(',', ',')# 获取标签tags = []tags_message = soup.select('.u-tag i')for p in tags_message:tags.append(p.get_text())# 对标签进行格式化if len(tags) > 1:tag = '-'.join(tags)else:tag = tags[0]# 获取歌单介绍if soup.select('#album-desc-more'):text = soup.select('#album-desc-more')[0].get_text().replace('\n', '').replace(',', ',')else:text = '无'# 获取歌单收藏量collection = soup.select('#content-operation i')[1].get_text().replace('(', '').replace(')', '')# 歌单播放量play = soup.select('.s-fc6')[0].get_text()# 歌单内歌曲数songs = soup.select('#playlist-track-count')[0].get_text()# 歌单评论数comments = soup.select('#cnt_comment_count')[0].get_text()# 输出歌单详情页信息print(title, tag, text, collection, play, songs, comments)# 将详情页信息写入CSV文件中with open('music_message.csv', 'a+', encoding='utf-8-sig') as f:f.write(title + ',' + tag + ',' + text + ',' + collection + ',' + play + ',' + songs + ',' + comments + '\n')# 获取歌单内歌曲名称li = soup.select('.f-hide li a')for j in li:with open('music_name.csv', 'a+', encoding='utf-8-sig') as f:f.write(j.get_text() + '\n')

最后获取的详情页如下:

本内容不代表本网观点和政治立场,如有侵犯你的权益请联系我们处理。
网友评论
网友评论仅供其表达个人看法,并不表明网站立场。