遍历链接数组会导致导航超时错误-puppeteer

如何解决遍历链接数组会导致导航超时错误-puppeteer

我有一个按钮元素数组,我想要一个按钮一次单击,并对每个打开的新选项卡执行此操作:

  1. 抓取一些信息并存储在名为“ providers”的数组中
  2. 关闭该标签页

虽然我能够执行此操作,但由于在browser.pages()之前使用了导航组件,所以我一直收到超时错误。如果删除组件,则会收到另一个超时错误。另外,每次我运行该程序时,在按钮数组进行不同次数的迭代后,都会遇到超时错误。这是我的代码:

const puppeteer = require("puppeteer");

(async () => {
  try {
    const browser = await puppeteer.launch({
      headless: false,});
    const page = await browser.newPage();

    //google.com
    await page.setExtraHTTPHeaders({ "Accept-Language": "en-US" });
    await page.goto("https://google.com");
    await page.type("input.gLFyf.gsfi","hotels in london");
    await page.keyboard.press("Enter");

    //search results
    await page.waitForXPath('//span[contains(text(),"View ")]');
    const btn1 = await page.$x('//span[contains(text(),"View ")]');
    await btn1[0].click();

    //list of hotels
    await page.waitForXPath('//span[contains(text(),"Learn more")]');

    let hotels = [];
    
    //buttons array that contains a list of buttons
    let buttons = await page.$x("//button[contains(.,'View prices')]");
 
    //prints a different value each time the program is run
    console.log(buttons.length);
 
    //looping through buttons array
    for (var i = 0; i < buttons.length; i++) {

      //i = 1 or 0 when program hangs 
      console.log("got here " + I);

      //*******************************cause of timeout error******************************************

      await page.setDefaultNavigationTimeout(0);
      await Promise.all([
        page.waitForNavigation({ waitUntil: "load",timeout: 0 }),buttons[i].click(),]);

      //***********************************************************************************************

      //getting all open tabs in an array
      const pages = await browser.pages();
      const page2 = pages[pages.length - 1];
      console.log(pages.length);

      //newly opened tab,sometimes program hangs before opening a new tab
      await page2
        .waitForSelector(
          "#prices > c-wiz > div > div.G86l0b > div > div > div > div > div > section > div.Hkwcrd.q9W60.A5WLXb.fLClSe > c-wiz > div > div > span > div > div > div > div > div > a > div > div.cFdfnb > div > span.mK0tQb > span",{ timeout: 30000 }
        )
        .catch(() => console.log("Class doesn't exist!"));

      /*-----------------scraping information on new tab ----------------------------------*/

      console.log("going to start collecting providers");
      let providers = await page2.evaluate(() => {
        let data = [];
        let elements = document.querySelectorAll(
          "#prices > c-wiz > div > div.G86l0b > div > div > div > div > div > section > div.Hkwcrd.q9W60.A5WLXb.fLClSe > c-wiz > div > div > span > div > div > div > div > div > a > div > div.cFdfnb > div > span.mK0tQb > span"
        );
        for (var element of elements) data.push(element.textContent);
        return data;
      });
      console.log(providers.length);
      console.log("all done");
      console.log(providers);
      hotels.push(providers);

      //closing the new tab
      page2.close();
    }
    
    await browser.close();
    return hotels;
  } catch (err) {
    console.error(err);
  }
})()
  .then((resolvedValue) => {
    console.log(resolvedValue);
  })
  .catch((rejectedValue) => {
    console.log(rejectedValue);
  });


要消除该错误,我使用了超时:0和setDefaultNavigationTimeout(0),但是现在程序冻结了。这是我在禁用超时获取之前遇到的错误:

TimeoutError: Navigation timeout of 30000 ms exceeded
    at C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\LifecycleWatcher.js:100:111
    at async FrameManager.waitForFrameNavigation (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\FrameManager.js:107:23)
    at async Frame.waitForNavigation (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\FrameManager.js:298:16)
    at async Page.waitForNavigation (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\Page.js:560:16)
    at async Promise.all (index 0)
    at async C:\Users\Me\Desktop\web_scraping_practice\backend.js:41:7
  -- ASYNC --
    at Frame.<anonymous> (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\helper.js:116:19)
    at Page.waitForNavigation (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\Page.js:560:53)
    at Page.<anonymous> (C:\Users\Me\Desktop\web_scraping_practice\node_modules\puppeteer\lib\helper.js:117:27)
    at C:\Users\Me\Desktop\web_scraping_practice\backend.js:42:14
    at processTicksAndRejections (internal/process/task_queues.js:97:5) {
  name: 'TimeoutError'
}
undefined

谢谢

解决方法

尝试运行您的代码,如果您按其内容搜索跨度,似乎最好对Chromium语言环境进行硬编码,因为在我的浏览器中,它们不是英语的。但是我做了一些调整,设法打开了一个包含酒店详细信息的标签。问题是这个选择器:

$("#prices > c-wiz > div > div.G86l0b > div > div > div > div > div > section > div.Hkwcrd.q9W60.A5WLXb.fLClSe > c-wiz > div > div > span > div > div > div > div > div > a > div > div.cFdfnb > div > span.mK0tQb > span");

不幸的是,这个东西渲染了null。我相信这组类div.Hkwcrd.q9W60.A5WLXb.fLClSe是动态生成的。不知道您实际上要提取什么信息,但是我会尝试通过此data-click-type属性来查找DOM元素。就我而言,它产生:

document.querySelectorAll("div[data-click-type='283']");
NodeList(18) [div.YPrvOd,div.YPrvOd,div.YPrvOd]

似乎是房间类型(高级双人间等)。 “ 268”点击类型似乎是包含酒店(预订,hotels.com等)的网站

下面的代码:

const puppeteer = require("puppeteer");

(async () => {
  try {
    const browser = await puppeteer.launch({
      headless: false,});
    const page = await browser.newPage();

    //google.com
    await page.setExtraHTTPHeaders({ "Accept-Language": "en-US" });
    await page.goto("https://google.com");
    await page.type("input.gLFyf.gsfi","hotels in london");
    await page.keyboard.press("Enter");

    //search results
    await page.waitForXPath('//span[contains(text(),"View ")]');
    const btn1 = await page.$x('//span[contains(text(),"View ")]');
    await btn1[0].click();

    //list of hotels
    await page.waitForXPath('//span[contains(text(),"Learn more")]');

    let hotels = [];

    //buttons array that contains a list of buttons
    let buttons = await page.$x("//button[contains(.,'View prices')]");

    //prints a different value each time the program is run
    console.log(buttons.length);

    //looping through buttons array
    for (var i = 0; i < buttons.length; i++) {

      //i = 1 or 0 when program hangs
      console.log("got here " + i);

      //*******************************cause of timeout error******************************************

      await page.setDefaultNavigationTimeout(0);
      await Promise.all([
        page.waitForNavigation({ waitUntil: "load",timeout: 0 }),buttons[i].click(),]);

      //***********************************************************************************************

      //getting all open tabs in an array
      const pages = await browser.pages();
      const page2 = pages[pages.length - 1];
      console.log(pages.length);

      //newly opened tab,sometimes program hangs before opening a new tab
      await page2
        .waitForSelector(
          "span[data-click-type='268']",{ timeout: 30000 }
        )
        .catch(() => console.log("Class doesn't exist!"));

      /*-----------------scraping information on new tab ----------------------------------*/

      console.log("going to start collecting providers");
      let providers = await page2.evaluate(() => {
        let data = [];
        let elements = document.querySelectorAll(
          "span[data-click-type='268']"
        );
        for (var element of elements) data.push(element.textContent);
        return data;
      });
      console.log(providers.length);
      console.log("all done");
      console.log(providers);
      hotels.push(providers);

      //closing the new tab
      page2.close();
    }

    await browser.close();
    return hotels;
  } catch (err) {
    console.error(err);
  }
})()
  .then((resolvedValue) => {
    console.log(resolvedValue);
  })
  .catch((rejectedValue) => {
    console.log(rejectedValue);
  });

在我的情况下,呈现以下内容:

(node:16816) ExperimentalWarning: The fs.promises API is experimental
12
got here 0
3
going to start collecting providers
16
all done
[ 'Booking.com','Tripadvisor.com','Agoda','Hotels.com','Booking.com','Expedia.com','Destinia','Stayforlong.com','Trip.com','ebookers.ie','Etrip','ZenHotels.com','Nustay.com' ]
got here 1

我相信,这是providers的列表。 通知使用的选择器:span[data-click-type='268']

版权声明:本文内容由互联网用户自发贡献,该文观点与技术仅代表作者本人。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如发现本站有涉嫌侵权/违法违规的内容, 请发送邮件至 dio@foxmail.com 举报,一经查实,本站将立刻删除。

相关推荐


依赖报错 idea导入项目后依赖报错,解决方案:https://blog.csdn.net/weixin_42420249/article/details/81191861 依赖版本报错:更换其他版本 无法下载依赖可参考:https://blog.csdn.net/weixin_42628809/a
错误1:代码生成器依赖和mybatis依赖冲突 启动项目时报错如下 2021-12-03 13:33:33.927 ERROR 7228 [ main] o.s.b.d.LoggingFailureAnalysisReporter : *************************** APPL
错误1:gradle项目控制台输出为乱码 # 解决方案:https://blog.csdn.net/weixin_43501566/article/details/112482302 # 在gradle-wrapper.properties 添加以下内容 org.gradle.jvmargs=-Df
错误还原:在查询的过程中,传入的workType为0时,该条件不起作用 &lt;select id=&quot;xxx&quot;&gt; SELECT di.id, di.name, di.work_type, di.updated... &lt;where&gt; &lt;if test=&qu
报错如下,gcc版本太低 ^ server.c:5346:31: 错误:‘struct redisServer’没有名为‘server_cpulist’的成员 redisSetCpuAffinity(server.server_cpulist); ^ server.c: 在函数‘hasActiveC
解决方案1 1、改项目中.idea/workspace.xml配置文件,增加dynamic.classpath参数 2、搜索PropertiesComponent,添加如下 &lt;property name=&quot;dynamic.classpath&quot; value=&quot;tru
删除根组件app.vue中的默认代码后报错:Module Error (from ./node_modules/eslint-loader/index.js): 解决方案:关闭ESlint代码检测,在项目根目录创建vue.config.js,在文件中添加 module.exports = { lin
查看spark默认的python版本 [root@master day27]# pyspark /home/software/spark-2.3.4-bin-hadoop2.7/conf/spark-env.sh: line 2: /usr/local/hadoop/bin/hadoop: No s
使用本地python环境可以成功执行 import pandas as pd import matplotlib.pyplot as plt # 设置字体 plt.rcParams[&#39;font.sans-serif&#39;] = [&#39;SimHei&#39;] # 能正确显示负号 p
错误1:Request method ‘DELETE‘ not supported 错误还原:controller层有一个接口,访问该接口时报错:Request method ‘DELETE‘ not supported 错误原因:没有接收到前端传入的参数,修改为如下 参考 错误2:cannot r
错误1:启动docker镜像时报错:Error response from daemon: driver failed programming external connectivity on endpoint quirky_allen 解决方法:重启docker -&gt; systemctl r
错误1:private field ‘xxx‘ is never assigned 按Altʾnter快捷键,选择第2项 参考:https://blog.csdn.net/shi_hong_fei_hei/article/details/88814070 错误2:启动时报错,不能找到主启动类 #
报错如下,通过源不能下载,最后警告pip需升级版本 Requirement already satisfied: pip in c:\users\ychen\appdata\local\programs\python\python310\lib\site-packages (22.0.4) Coll
错误1:maven打包报错 错误还原:使用maven打包项目时报错如下 [ERROR] Failed to execute goal org.apache.maven.plugins:maven-resources-plugin:3.2.0:resources (default-resources)
错误1:服务调用时报错 服务消费者模块assess通过openFeign调用服务提供者模块hires 如下为服务提供者模块hires的控制层接口 @RestController @RequestMapping(&quot;/hires&quot;) public class FeignControl
错误1:运行项目后报如下错误 解决方案 报错2:Failed to execute goal org.apache.maven.plugins:maven-compiler-plugin:3.8.1:compile (default-compile) on project sb 解决方案:在pom.
参考 错误原因 过滤器或拦截器在生效时,redisTemplate还没有注入 解决方案:在注入容器时就生效 @Component //项目运行时就注入Spring容器 public class RedisBean { @Resource private RedisTemplate&lt;String
使用vite构建项目报错 C:\Users\ychen\work&gt;npm init @vitejs/app @vitejs/create-app is deprecated, use npm init vite instead C:\Users\ychen\AppData\Local\npm-