-
Notifications
You must be signed in to change notification settings - Fork 136
/
Copy pathREADME.Rmd
149 lines (112 loc) · 6.25 KB
/
README.Rmd
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
---
output: github_document
---
```{r setup, include=FALSE}
knitr::opts_chunk$set(echo = TRUE, collapse = TRUE, comment = "#>", eval = FALSE)
```
# tweetbotornot <img width="160px" src="man/figures/logo.png" align="right" />
[![lifecycle](https://img.shields.io/badge/lifecycle-experimental-orange.svg)](https://www.tidyverse.org/lifecycle/#experimental)
[![Travis build status](https://travis-ci.org/mkearney/tweetbotornot.svg?branch=master)](https://travis-ci.org/mkearney/tweetbotornot)
[![Coverage status](https://codecov.io/gh/mkearney/tweetbotornot/branch/master/graph/badge.svg)](https://codecov.io/github/mkearney/tweetbotornot?branch=master)
An R package for classifying Twitter accounts as `bot or not`.
## Features
Uses machine learning to classify Twitter accounts as bots or not bots. The **default model** is 93.53% accurate when classifying bots and 95.32% accurate when classifying non-bots. The **fast model** is 91.78% accurate when classifying bots and 92.61% accurate when classifying non-bots.
Overall, the **default model** is correct 93.8% of the time.
Overall, the **fast model** is correct 91.9% of the time.
## Install
Install from CRAN:
``` r
## install from CRAN
install.packages("tweetbotornot")
```
Install the development version from Github:
``` r
## install remotes if not already
if (!requireNamespace("remotes", quietly = TRUE)) {
install.packages("remotes")
}
## install tweetbotornot from github
devtools::install_github("mkearney/tweetbotornot")
```
## API authorization
Users must be authorized in order to interact with Twitter's API. To setup your
machine to make authorized requests, you'll either need to be signed into Twitter and working in an interactive session of R–the browser will open asking you to authorize the rtweet client (rstats2twitter)–or you'll need to create an app (and have a developer account) and your own API token. The latter has the benefit of (a) having sufficient permissions for write-acess and DM (direct messages) read-access levels and (b) more stability if Twitter decides to shut down [@kearneymw](https://twitter.com/kearneymw)'s access to Twitter (I try to be very responsible these days, but Twitter isn't always friendly to academic use cases). To create an app and your own Twitter token, [see these instructions provided in the rtweet package](http://rtweet.info/articles/auth.html).
## Usage
There's one function `tweetbotornot()` (technically there's also `botornot()`, but it does the same exact thing).
Give it a vector of screen names or user IDs and let it go to work.
```{r}
## load package
library(tweetbotornot)
## select users
users <- c("realdonaldtrump", "netflix_bot",
"kearneymw", "dataandme", "hadleywickham",
"ma_salmon", "juliasilge", "tidyversetweets",
"American__Voter", "mothgenerator", "hrbrmstr")
## get botornot estimates
data <- tweetbotornot(users)
## arrange by prob ests
data[order(data$prob_bot), ]
#> # A tibble: 11 x 3
#> screen_name user_id prob_bot
#> <chr> <chr> <dbl>
#> 1 hadleywickham 69133574 0.00754
#> 2 realDonaldTrump 25073877 0.00995
#> 3 kearneymw 2973406683 0.0607
#> 4 ma_salmon 2865404679 0.150
#> 5 juliasilge 13074042 0.162
#> 6 dataandme 3230388598 0.227
#> 7 hrbrmstr 5685812 0.320
#> 8 netflix_bot 1203840834 0.978
#> 9 tidyversetweets 935569091678691328 0.997
#> 10 mothgenerator 3277928935 0.998
#> 11 American__Voter 829792389925597184 1.000
```
### Integration with rtweet
The `botornot()` function also accepts data returned by [rtweet](http://rtweet.info) functions.
```{r}
## get most recent 100 tweets from each user
tmls <- get_timelines(users, n = 100)
## pass the returned data to botornot()
data <- botornot(tmls)
## arrange by prob ests
data[order(data$prob_bot), ]
#> # A tibble: 11 x 3
#> screen_name user_id prob_bot
#> <chr> <chr> <dbl>
#> 1 hadleywickham 69133574 0.00754
#> 2 realDonaldTrump 25073877 0.00995
#> 3 kearneymw 2973406683 0.0607
#> 4 ma_salmon 2865404679 0.150
#> 5 juliasilge 13074042 0.162
#> 6 dataandme 3230388598 0.227
#> 7 hrbrmstr 5685812 0.320
#> 8 netflix_bot 1203840834 0.978
#> 9 tidyversetweets 935569091678691328 0.997
#> 10 mothgenerator 3277928935 0.998
#> 11 American__Voter 829792389925597184 1.000
```
### `fast = TRUE`
The default [gradient boosted] model uses both users-level (bio, location, number of followers and friends, etc.) **and** tweets-level (number of hashtags, mentions, capital letters, etc. in a user's most recent 100 tweets) data to estimate the probability that users are bots. For larger data sets, this method can be quite slow. Due to Twitter's REST API rate limits, users are limited to only 180 estimates per every 15 minutes.
To maximize the number of estimates per 15 minutes (at the cost of being less accurate), use the `fast = TRUE` argument. This method uses **only** users-level data, which increases the maximum number of estimates per 15 minutes to *90,000*! Due to losses in accuracy, this method should be used with caution!
```{r}
## get botornot estimates
data <- botornot(users, fast = TRUE)
## arrange by prob ests
data[order(data$prob_bot), ]
#> # A tibble: 11 x 3
#> screen_name user_id prob_bot
#> <chr> <chr> <dbl>
#> 1 hadleywickham 69133574 0.00185
#> 2 kearneymw 2973406683 0.0415
#> 3 ma_salmon 2865404679 0.0661
#> 4 dataandme 3230388598 0.0965
#> 5 juliasilge 13074042 0.112
#> 6 hrbrmstr 5685812 0.121
#> 7 realDonaldTrump 25073877 0.368
#> 8 netflix_bot 1203840834 0.978
#> 9 tidyversetweets 935569091678691328 0.998
#> 10 mothgenerator 3277928935 0.999
#> 11 American__Voter 829792389925597184 0.999
```
## NOTE
In order to avoid confusion, the package was renamed from "botrnot" to "tweetbotornot" in June 2018. This package should not be confused with the [botornot application](http://botornot.co/).