Uses This

UsesThis.com has >1000 interviews answering the same 4 questions related to how people work.

Skip to the plots, starting at #OS Usage

Outside Inspiration

I'm trying to find setups that are outside. There are lots of matches, especially in dream setups. Some other useful search terms: sun, grass, porch

grep -P '\b(sun|weather|outside|porch)\b'  usesthis_pages/* |grep -Pv "outside (of|that|your)" | pandoc -f html -t plain
grep -P '\b(grass)'  usesthis_pages/* | grep -v Grasshopper | pandoc -f html -t plain

Ultimately, I didn't get an additional insight for my own setup. But did confirm many like working outside.

I also wish I'd symlinked the filenames to include date and the author categories. Then grep could have had some useful metatdata. As it is, I just jumped to the matching file.

More details

Now that there's 5Mb of interviews locally, I want to poke at it a bit.

Extracting

UsesThis links tools and hardware (with useful title attributes too). They can be pulled out like

<a[^>]+(?<title>title="[^"]+)[^>]*>(?<name>.*?)<\/a

But there are a few links on every page. I try to exclude a those

href="(/"|.*ko-fi.com|/about|/categories/")|Attribution-Share Alike|subscribing to the feed

Tool links occur anywhere in the 4 Questions asked:

  • Who are you, and what do you do?
  • What would be your dream setup?
  • What hardware do you use (first 14 are What do you use to get the job done?)
  • And what software?

I pull these into $sec (section; would be better named as question)

\QWho are you, and what do you do?\E|
\QWhat would be your dream setup?\E |
\QWhat hardware do you use\E        |
\QAnd what software?\E              |
\QWhat do you use to get the job done?\E

Publish time is also well annotated.

time datetime="([^"]+)

Extracting all in one pass with perl. This started as perl -lne 'BEGIN{} ...' usesthis_pages/* > usesthis_tags.txt, but to explore org mode/literate programming. I moved into noweb and perl.

open STDOUT, '>', "usesthis_tagsx.txt"; # bash '> usethis_tags.txt'
@ARGV = glob('usesthis_pages/*');       # 'perl -n' part 1
sub resetf {$dt=0; $sec="top";}         # first section and time
resetf;
while(<>){                              # 'perl -n' part 2
  $dt=$1 if m/<<regex-date>>/;
  $sec=$& if m/<<regex-sec>>/x;
  while(/<<regex-tag>>/g){
   $t=$+{title};
   $n=$+{name};
   next if $& =~ m:<<regex-exclude>>:;
   print("$dt\t",
         $ARGV=~s:.*/::r ,
         "\t$sec\t$n\t",
         $t=~s/.*?"|\.$//gr, "\n");
  }
  resetf if(eof);
}

Org-mode org-babel noweb aside

Originally, the perl-to-long-tag extraction was iterated on in a terminal using an expansive perl one-liner. Once it was running as expected, I run on the full glob (all files). Then removed the bash wildcard with perl glob() and > redirect with open STDOUT I also moved the regular expressions into their own code blocks and weaved them in with noweb (eg. <<regex-date>>).

But arrant angles << >> messed with perl syntax highlighting. I moved the pattern delimiters and modifiers m:...:x back into the code. This means that there's context need to understand the regexp that's only in the code:

  • x modifier for ignore spacing
  • m:...: colon deliminator so slash doesn't have to be escaped

What's more :noweb-prefix no is needed in the babel headers to correctly use the multi-line pattern.

Tool use

pacman::p_load(dplyr,readr,stringr, ggplot2)
clps <- \(x) x |> str_trunc(width=20) |> sort() |> unique() |>  paste0(collapse=";")

long <- read_delim("usesthis_tags.txt",
                   col_names=c("published", "who","sec","name","desc"),
                   delim="\t") |>
  mutate(published=lubridate::ymd(published)) |>
  filter(!is.na(published)) # 56 for whom name was before datetime

tags <- long |>
  group_by(name)|>
  summarise(n=n(), fdate=min(published), ldate=max(published),
            desc=clps(desc)) |>
  arrange(-n)

tags|>head()
namenfdateldatedesc
mac9482009-01-022026-09-26All the mac inter…
developer4912009-01-022026-09-26All the developer…;Charlie&#39;s Git…;Mitchell&#39;s Gi…
MacBook Pro4362009-01-022026-08-13A laptop
Photoshop3782009-01-092026-04-30A bitmap image ed…
windows3092009-01-172026-09-05All the windows i…
Chrome2662009-01-172025-05-30A web browser;A web browser app;A WebKit-based br…

OS

"top" section (before questions) contains category tags. These include operating system and biographical info ("developer", "artist"). Interviewees can and do endorse more than one of either.

os <- long |>
  filter(name %in% c('linux','mac','windows')) |>
  group_by(who) |>
  summarise(published=first(published),
	    os_long=clps(name),
            os=ifelse(grepl(';',os_long), "mutli", first(name)))

os |> count(os_long) |> arrange(-n)
os_longn
mac763
windows166
mac;windows93
linux64
linux;mac64
linux;mac;windows28
linux;windows22

To simplify, each interviewee is assigned the category they endorsed that is most often endorsed by others: artist, developer becomes just developer.

cats <- long |>
  filter(sec=="top", ! name %in% c('mac','linux','windows')) |>
  count(name)

main_desc <- cats |> inner_join(long) |>
  group_by(who) |> filter(n==max(n)) |> ungroup()

category_cnt <- main_desc |> count(name) |> arrange(-n)
head(category_cnt)
namen
developer491
writer212
designer158
artist76
musician60
game42

Who uses linux?

cat_os <- main_desc |> inner_join(os)

# user of linux (alone or w/mac and or win)
# 'name' is catagory (dev, lawyer, hacker, etc)
cat_os |> select(-n) |>
  filter(grepl('linux',os_long)) |>
  count(name, name='linux') |>
  # merge to catagory counts (across all endoursments, many per person)
  inner_join(cats)  |>
  # calc percent
  transmute(name,linux, total=n, percent=round(linux/total*100,2)) |>
  arrange(-percent) |>
  head(n=10)
namelinuxtotalpercent
developer13948928.43
service2825
hacker83324.24
lawyer1520
politics1520
professor53315.15
scientist53315.15
engineer1119.09
hardware1119.09
researcher3358.57

OS Usage

p_os <- long |>
  filter(name %in% c('linux','mac','windows')) |>
ggplot() + aes(x=published, fill=name) +
  geom_histogram(position="stack",bins=10) +
  theme_bw() + facet_wrap(~name) + labs(title="OS by interview", fill='os')

p_cat <- cat_os |>
  filter(name %in% category_cnt$name[1:5]) |>
 ggplot() + aes(x=published, fill=os) + geom_histogram() +
 theme_bw() + facet_wrap(~name)

pacman::p_load(patchwork)
(p_os+guides(fill="none")+labs(x="")) / p_cat +
  plot_layout(guides = "collect", axis="collect", heights = c(1, 2))

../images/usesthis/os.png

Software

I'm glad firefox isn't lagging too far behind chrome. And emacs and vim are top editors.

tools <- long |>
  filter(grepl('what software|get the job', sec)) |>
  transmute(who,
            tool=tolower(name) |> gsub('[^a-z0-9]','',x=_)|>
	      gsub('googlechrome','chrome',x=_)|>
	      gsub('mozillafirefox','firefox',x=_)|>
	      gsub('sublimet.*','sublime',x=_)|>
	      gsub('gnuemacs','emacs',x=_) |>
	      gsub('i3tile.*','i3',x=_) |>
	      gsub('rhymeswithgonad','xmonad',x=_) |>
	      gsub('visualstudiocode','vscode',x=_) |>
	      gsub('intellijidea','intellij',x=_) |>
	      gsub('^vi$','vim',x=_) |>
	      gsub('awesomew.*','awesome',x=_),
	    desc_cat=desc|> tolower() |>
	      gsub('irc ','chat',x=_)|>
	      gsub('ide\\b','text editor ',x=_)|>
	      gsub('terminal shell','shell',x=_)|>
              str_extract('chat|browser|window manager|text editor|term|shelll')) |>
  merge(cat_os, by='who')

linux_tools <- tools |>
  filter(grepl('linux',os_long)) |>
  count(tool, name, desc_cat) |>
  arrange(-n)

tools_totalsum <- linux_tools |>
  group_by(tool) |>
  summarise(nsum=sum(n)) |>
  merge(linux_tools, by="tool") |>
  filter(!is.na(desc_cat),nsum>2)

ggplot(tools_totalsum) +
 aes(x=tool, fill=name, y=n) +
  geom_histogram(stat='identity') +
  facet_wrap(~desc_cat, scales="free_x") +
  see::theme_modern(axis.text.angle=45)

../images/usesthis/tools.png

Editors

Over time we can see vscode take off, emacs ebb and flow, atom pop in and out, eclipse fade away.

all_tools <- tools |> filter(desc_cat %in% c('text editor'))

all_tools |> count(tool,name='ntool') |>
  merge(all_tools, by='tool') |>
  filter(tool %in% tools_totalsum$tool) |>
  #filter(ntool>2) |>
ggplot() + aes(x=published, y=tool, color=os) + geom_point() +
facet_grid(desc_cat~., scales="free") + theme_bw() +
labs(title="Text editors over time")

../images/usesthis/editor.png

LLM inference usage

  • prompted chatgpt for skip-existing option in curl --remote-name-all, circled around --no-clobber to decided that wouldn't work (makes *.1 instead of overwrite)
  • perl -lne '...' gl*ob > out to @ARGV=...; open STDOUT, ... was a mix of chatGPT hallucination, Claude, stack overflow, and perldoc.
  • uploaded pandoc converted plaintext file of all to query with sonnet 5 medium. prompt for summary, specific hardware used outside. confirmed what was found with grep.

..