I am building a suite of tools to automate Twitter functions outside the paid API.

Bookmarks Scraper (July 2024) Github

I wanted to download all my bookmarked images and posts from Twitter and index them, but it costed $100/mo to do this with the official Twitter API. So, I built a cost-effective workaround.

The Method

  1. Scraping the data
  2. Analyzing the data

Areas for Expansion and Improvement

This scraper is an ongoing project with potential research-level scaleability (as the paid API effectively limits a lot of researchers). Improvements include:

Previous Iterations

  1. The initial attempt was to use selenium to automate scrolling and BeautifulSoup to parse the resulting HTML. This was slow, as all the images had to load each time; painful, as using BeautifulSoup to parse HTML sucks; and fragile, as it would break if twitter changed their site structure.
  2. A next attempt was to use tshark (terminal wireshark) to analyze the network traffic. Decoding encrypted https (443) traffic with cookies was a pain and I could only get it working on Windows. Curl handled this automatically for me so I switched to that.

Mass Unfollow Dashboard (Oct 2024)

I was following too many accounts on my facename twitter account so I made a dashboard to quickly go through them all and unfollow them.

  1. Get all the accounts I follow using curl using this script
  2. Put them all into an html document in table format using this script. It shows relevant information and has checkbox columns for unfollowing and for putting into categories, and exports that information into json.
  3. This script uses selenium to go through the json and add all the accounts to lists and unfollow them.