Skip to main content

Webscraping with RStudio

I'm trying to webscrap a website (actually I'm downloading a .xlsx file) with R, but I got stuck since I don't know too much about both R and HTML.

I found the code below here, but it is not doing all the job. On the website I need to select (on a dropdown menu) the department and the checklist type, and then click on a download button (which opens the "download url" from the code bellow). The problem is the website allows me to download data from one department at a time.

So I would like to find a way to do this with R, and download the files from all the departments just changing the ID (from both department and checklist).

install.packages("rvest")
library(rvest)

url <- "https://website.com/auth/login?redirectUrl="
download_url <- "https://website.com/departments/file.xlsx"
session <- html_session(url)
form <- html_form(session)[[1]]

filled_form <- set_values(form,
                          email = "myusername", # "email" is an element from website HTML code
                          password = "mypassword") # "mypassword" is an element from website HTML code

## Save main page url
main_page <- submit_form(session, filled_form)

#after login and submit_form do this:
download <- jump_to(main_page, download_url)

# write file to current working directory
writeBin(download$response$content, basename(download_url))

The area ID has the format in the HTML code bellow. I guess the JS function gets this setIdPlantaAtual(123) value and filter the database when I select the download button. The other filter seems to be based on the setModuloAtual parameter (equivalent to checklist type).

// Department HTML code

<h5><a onclick="setIdPlantaAtual(123)" href="javascript:;">Department ABC <span class="badge badge-success">selecionar</span></a></h5>

// Checklist HTML code

<li class="new" onclick="setModuloAtual('Checklist1')" style="cursor:pointer;">

Bellow are the two functions I found on the website code.

        function setIdPlantaAtual(_idplantaatual) {
            $.blockUI({
                message: '<h1> Wait...</h1>'
            });
            $.post(baseUrl + '/index/idplantaatual', {
                id: _idplantaatual
            }, function(res) {
                if (res == 'OK') {
                    location.reload();
                } else {
                    $.unblockUI();
                }
            });

        }

        function setModuloAtual(modulo) {
            if (modulo == "Tratativas de NCs") {
                modulo = "Tratativas de NC\'s";
            } 
           
            $.post(baseUrl + '/index/moduloatual', {
                modulo: modulo
            }, function(res) {
                console.log('aaa',res)
                if (res == 'OK') {
                    // location.reload();
                    window.location.replace(baseUrl);
                } else if (res.split('$%$').length > 0) {
                    var new_modulo = res.split('$%$')[1];
                    if (new_modulo != undefined) {
                        window.location.replace(baseUrl + '/' + new_modulo);
                    } else {
                        window.location.replace(baseUrl + '/');
                    }
                } else {
                    console.log(res);
                }
            });
        }

I think the same approach from my first attempt would work to merge the department ID (setIdPlantaAtual()) and the checklist ID (setModuloAtual()) into the download url, but I don't know how to do that.

I appreciate any help!

Via Active questions tagged javascript - Stack Overflow https://ift.tt/Wiv3CQa

Comments

Popular posts from this blog

How to show number of registered users in Laravel based on usertype?

i'm trying to display data from the database in the admin dashboard i used this: <?php use Illuminate\Support\Facades\DB; $users = DB::table('users')->count(); echo $users; ?> and i have successfully get the correct data from the database but what if i want to display a specific data for example in this user table there is "usertype" that specify if the user is normal user or admin i want to user the same code above but to display a specific usertype i tried this: <?php use Illuminate\Support\Facades\DB; $users = DB::table('users')->count()->WHERE usertype =admin; echo $users; ?> but it didn't work, what am i doing wrong? source https://stackoverflow.com/questions/68199726/how-to-show-number-of-registered-users-in-laravel-based-on-usertype

Why is my reports service not connecting?

I am trying to pull some data from a Postgres database using Node.js and node-postures but I can't figure out why my service isn't connecting. my routes/index.js file: const express = require('express'); const router = express.Router(); const ordersCountController = require('../controllers/ordersCountController'); const ordersController = require('../controllers/ordersController'); const weeklyReportsController = require('../controllers/weeklyReportsController'); router.get('/orders_count', ordersCountController); router.get('/orders', ordersController); router.get('/weekly_reports', weeklyReportsController); module.exports = router; My controllers/weeklyReportsController.js file: const weeklyReportsService = require('../services/weeklyReportsService'); const weeklyReportsController = async (req, res) => { try { const data = await weeklyReportsService; res.json({data}) console...

How to split a rinex file if I need 24 hours data

Trying to divide rinex file using the command gfzrnx but getting this error. While doing that getting this error msg 'gfzrnx' is not recognized as an internal or external command Trying to split rinex file using the command gfzrnx. also install'gfzrnx'. my doubt is I need to run this program in 'gfzrnx' or in 'cmdprompt'. I am expecting a rinex file with 24 hrs or 1 day data.I Have 48 hrs data in RINEX format. Please help me to solve this issue. source https://stackoverflow.com/questions/75385367/how-to-split-a-rinex-file-if-i-need-24-hours-data