Work Motivation Under Lockdown Exploratory Data Analysis (EDA)

Two datasets...

  1. Post-workday evaluations of work motivation related stuff. Large spread across how many time points people collected; few decent-ish ones.

  2. Traditional questionnaires from the same participants.

    • Allocation indicates whether they took part in our motivation self-management training during the period or belonged to a waitlist control group.
    • Pre-, mid- and post-questionnaires --> TIME == 1, 2 and 3 respectively.
    • Everything from intrinsic1 to Psycap5 are raw answers. From Intrinsic to Optimism are sumscores.
    • This was a feasibility & acceptability study so it's not powered for or aiming to look at outcomes. But if there was one, it could be the "Relative Autonomy Index" (RAI), a sumscore conventionally built as:
      • Amotivation -3 + ExternalTotal -2 + Introjected -1 + Identified 2 + Intrinsic * 3
    • If I remember correctly, there were some people whose RAI was boosted a lot
    • In general:
      • Good things to go up during the intervention (or life in general):
        • Intrinsic, Identified
        • Competence, Relatedness, Autonomy
      • Bad things to go up during the intervention (or life in general):
        • Amotivation, External (material & social), Introjected
        • AutonomyThwarting, RelatednessThwarting, CompetenceThwarting
          • Unsure if these were named in the data, or just some of the items named autonomyx/relatednessx/competencex

This EDA only addresses the daily data, specifically data=="data/moti_feasibility_james_daily.csv"

Show the session information of the packages used in this analysis

Read in the raw daily data

Reshape the data frame and convert data types to numeric where possible

Count the amount of rows per user and plot

The aim of this is to assess whether there is enough data to do some of the more complext NLTSA methods

Plot the heatmaps of the top 5 users with the highest amount of rows

The aim of this is to assess the amount of missingness time wise (i.e. the user could have a lot of data, spread over a long period with large gaps in the middle).

Visualize outliers

Check the normality using the Kolmogrov-Smirnov test

Run the correlations as sample size increases for the highest correlations

Level of aggreement between two clinical measurements accounting for multiple data points

Specify the model with the formula_string as the model syntax

Create a train and test set, use a fitted model to predict previously unseen rows (this is not controlled for user)

Users included in the training set

Users included in the test set

Users in the test set that were not included in the training set