动态网站爬虫Python-selenium-PhantomJS
Posted
tags:
篇首语:本文由小常识网(cha138.com)小编为大家整理,主要介绍了动态网站爬虫Python-selenium-PhantomJS相关的知识,希望对你有一定的参考价值。
from selenium import webdriver #from selenium.webdriver.common.proxy import Proxy from selenium.webdriver.common.proxy import ProxyType from selenium.webdriver.common.desired_capabilities import DesiredCapabilities dcap = dict(DesiredCapabilities.PHANTOMJS) dcap["phantomjs.page.settings.userAgent"] = ( "Mozilla/5.0 (iPod; U; CPU iPhone OS 2_1 like Mac OS X; ja-jp) AppleWebKit/525.18.1 (KHTML, like Gecko) Version/3.1.1 Mobile/5F137 Safari/525.20" )###设置浏览器hearders obj = webdriver.PhantomJS(executable_path="E:/phantomjs/bin/phantomjs.exe",desired_capabilities=dcap) proxy=webdriver.Proxy() proxy.proxy_type=ProxyType.MANUAL proxy.http_proxy=‘115.211.50.97:8118‘ proxy.add_to_capabilities(webdriver.DesiredCapabilities.PHANTOMJS) obj.start_session(webdriver.DesiredCapabilities.PHANTOMJS) obj.maximize_window() obj.get(‘http://fz12345.fuzhou.gov.cn/callcenter/callSearchDo.do‘) obj.set_page_load_timeout(15) #等待时间 obj.find_element_by_xpath("/html/body/div[2]/div/div[1]/div[5]/div[2]/div[2]/form/div[3]/ul/li[5]/a").click() a=obj.find_element_by_xpath("/html/body/div[2]/div/div[1]/div[5]/div[2]/div[2]/form/div[2]/table/tbody/tr[3]").text print(a) obj.close()
以上是关于动态网站爬虫Python-selenium-PhantomJS的主要内容,如果未能解决你的问题,请参考以下文章
爬虫之动态HTML处理(Selenium与PhantomJS )网站模拟登录