Files
matti_jms_collabs/work_motivation_EDA.ipynb
T

2.3 MiB

Work Motivation Under Lockdown Exploratory Data Analysis (EDA)

Two datasets...

  1. Post-workday evaluations of work motivation related stuff. Large spread across how many time points people collected; few decent-ish ones.

  2. Traditional questionnaires from the same participants.

    • Allocation indicates whether they took part in our motivation self-management training during the period or belonged to a waitlist control group.
    • Pre-, mid- and post-questionnaires --> TIME == 1, 2 and 3 respectively.
    • Everything from intrinsic1 to Psycap5 are raw answers. From Intrinsic to Optimism are sumscores.
    • This was a feasibility & acceptability study so it's not powered for or aiming to look at outcomes. But if there was one, it could be the "Relative Autonomy Index" (RAI), a sumscore conventionally built as:
      • Amotivation * -3 + ExternalTotal * -2 + Introjected * -1 + Identified * 2 + Intrinsic * 3
    • If I remember correctly, there were some people whose RAI was boosted a lot
    • In general:
      • Good things to go up during the intervention (or life in general):
        • Intrinsic, Identified
        • Competence, Relatedness, Autonomy
      • Bad things to go up during the intervention (or life in general):
        • Amotivation, External (material & social), Introjected
        • AutonomyThwarting, RelatednessThwarting, CompetenceThwarting
          • Unsure if these were named in the data, or just some of the items named autonomyx/relatednessx/competencex

This EDA only addresses the daily data, specifically data=="data/moti_feasibility_james_daily.csv"

In [1]:
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from matplotlib.colors import LinearSegmentedColormap
import session_info

from jmspack.NLTSA import (flatten, 
                          ts_levels,
                          fluctuation_intensity,
                          distribution_uniformity,
                          complexity_resonance,
                          complexity_resonance_diagram)

from sklearn.preprocessing import MinMaxScaler
# To use this experimental feature, we need to explicitly ask for it:
from sklearn.experimental import enable_iterative_imputer  # noqa
from sklearn.impute import IterativeImputer
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import ExtraTreesRegressor
from sklearn.neighbors import KNeighborsRegressor
from sklearn.pipeline import make_pipeline

from pyunicorn.timeseries import RecurrencePlot
pyunicorn: Package netCDF4 could not be loaded. Some functionality in class Data might not be available!
pyunicorn: Package netCDF4 could not be loaded. Some functionality in class NetCDFDictionary might not be available!

Show the session information of the packages used in this analysis

In [2]:
session_info.show(write_req_file=False,
                  req_file_name="work_motivation_EDA_requirements.txt",)
Out[2]:
Click to view session information
-----
jmspack             0.0.3
matplotlib          3.3.4
numpy               1.19.2
pandas              1.2.3
pyunicorn           NA
seaborn             0.11.1
session_info        1.0.0
sklearn             0.24.1
-----
Click to view modules imported as dependencies
PIL                 8.1.2
appnope             0.1.2
backcall            0.2.0
cffi                1.14.5
colorama            0.4.4
cycler              0.10.0
cython_runtime      NA
dateutil            2.8.1
decorator           4.4.2
igraph              0.9.1
ipykernel           5.3.4
ipython_genutils    0.2.0
ipywidgets          7.6.3
jedi                0.17.2
joblib              0.17.0
kiwisolver          1.3.1
mpl_toolkits        NA
parso               0.7.0
pexpect             4.8.0
pickleshare         0.7.5
pkg_resources       NA
prompt_toolkit      3.0.8
ptyprocess          0.7.0
pyexpat             NA
pygments            2.8.1
pyparsing           2.4.7
pytz                2021.1
scipy               1.6.1
six                 1.15.0
statsmodels         0.12.2
storemagic          NA
texttable           1.6.3
tornado             6.1
traitlets           5.0.5
wcwidth             0.2.5
zmq                 20.0.0
-----
IPython             7.21.0
jupyter_client      6.1.7
jupyter_core        4.7.1
jupyterlab          2.2.6
notebook            6.2.0
-----
Python 3.9.2 (default, Mar  3 2021, 11:58:52) [Clang 10.0.0 ]
macOS-10.16-x86_64-i386-64bit
-----
Session information updated at 2021-06-10 12:27

Read in the raw daily data

In [3]:
_df = pd.read_csv("data/moti_feasibility_james_daily.csv")
In [4]:
_df = _df.assign(date=pd.to_datetime(_df["Date (submitted)"].str.split("T", expand=True)[0], format="%Y-%m-%dT%H:%M:%S"))
In [5]:
def convert_numeric_where_possible(x):
    try:
        return x.astype(np.number)
    except:
        return x

Reshape the data frame and convert data types to numeric where possible

In [6]:
df = (_df
 .loc[:, ["User", "Field", "Value", "date"]]
 .pivot(index=["User", "date"], columns="Field")
 .droplevel(level=0, axis=1)
      .apply(convert_numeric_where_possible)
)
df.head(2)
Out[6]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
Field DailySituation absorption amotivation autonomy competence dedication emotional_drain enjoyment external pressure howami_participation_ended howami_participation_started importance interest internal pressure productivity_work relatedness satisfaction_work strategies time_pressure vigor
User date
Moti105 2020-03-03 NaN NaN NaN NaN NaN NaN NaN NaN NaN NaN 2020-03-03 NaN NaN NaN NaN NaN NaN NaN NaN NaN
2020-03-09 NaN 68.0 66.0 33.0 62.0 69.0 NaN 31.0 52.0 NaN NaN 67.0 74.0 60.0 72.0 65.0 42.0 emphasize_autonomy;reflect_desire;identifyas_r... NaN 64.0
In [7]:
df.info()
<class 'pandas.core.frame.DataFrame'>
MultiIndex: 1674 entries, ('Moti105', Timestamp('2020-03-03 00:00:00')) to ('Moti218', Timestamp('2020-05-15 00:00:00'))
Data columns (total 20 columns):
 #   Column                        Non-Null Count  Dtype  
---  ------                        --------------  -----  
 0   DailySituation                1173 non-null   object 
 1   absorption                    1487 non-null   float64
 2   amotivation                   1487 non-null   float64
 3   autonomy                      1487 non-null   float64
 4   competence                    1487 non-null   float64
 5   dedication                    1487 non-null   float64
 6   emotional_drain               1126 non-null   float64
 7   enjoyment                     1487 non-null   float64
 8   external pressure             1487 non-null   float64
 9   howami_participation_ended    11 non-null     object 
 10  howami_participation_started  101 non-null    object 
 11  importance                    1487 non-null   float64
 12  interest                      1487 non-null   float64
 13  internal pressure             1487 non-null   float64
 14  productivity_work             1487 non-null   float64
 15  relatedness                   1487 non-null   float64
 16  satisfaction_work             1487 non-null   float64
 17  strategies                    1301 non-null   object 
 18  time_pressure                 1126 non-null   float64
 19  vigor                         1487 non-null   float64
dtypes: float64(16), object(4)
memory usage: 274.8+ KB

Count the amount of rows per user and plot

The aim of this is to assess whether there is enough data to do some of the more complext NLTSA methods

In [8]:
row_amount_df = (df
.reset_index()
.groupby("User")
.count()
.loc[:, ["date"]]
.rename(columns={"date": "row_amount"})
.sort_values(by="row_amount")
.reset_index())
In [9]:
row_amount_df.tail(1)
Out[9]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
Field User row_amount
101 Moti148 64
In [10]:
_ = plt.figure(figsize=(20, 4))
_ = sns.barplot(data=row_amount_df, x="User", y="row_amount")
_ = plt.xticks(rotation=90)
_ = plt.axhline(30, c="red", ls="--", label="30 day cutoff")
_ = plt.title("Row amounts per user")
_ = plt.legend()
No description has been provided for this image

Plot the heatmaps of the top 5 users with the highest amount of rows

The aim of this is to assess the amount of missingness time wise (i.e. the user could have a lot of data, spread over a long period with large gaps in the middle).

In [11]:
top_users = row_amount_df.tail(5).User.tolist()
In [25]:
scale_data = True
for user in top_users:
    plot_df = df.select_dtypes(np.number).loc[(user, ), :]
    # plot_df.index.min(), plot_df.index.max()
    new_index = pd.date_range(start=plot_df.index.min(), end=plot_df.index.max())
    plot_df = plot_df.reindex(new_index)
    
    if scale_data:
        plot_df = pd.DataFrame(MinMaxScaler().fit_transform(plot_df), index=plot_df.index, columns=plot_df.columns)
    
    plot_df.index = plot_df.index.astype(str)
    
    _ = plt.figure(figsize=(20, 5))
    _ = sns.lineplot(data = plot_df.reset_index().melt(id_vars="index"), x="index", y="value", hue="Field")
    _ = sns.scatterplot(data = plot_df.reset_index().melt(id_vars="index"), x="index", y="value", hue="Field", legend=False)
    _ = plt.xticks(rotation=90)
    _ = plt.title(f"Lineplot of raw values, user == {user}")
    
    _ = plt.figure(figsize=(20, 5))
    _ = sns.heatmap(plot_df.T)
    _ = plt.title(f"Heatmap of raw values, user == {user}")
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image

Show the heatmap of the missing values for the last user plotted above

In [26]:
_ = plt.figure(figsize=(20, 5))
_ = sns.heatmap(plot_df.isna().astype(int).T)
_ = plt.title(f"Heatmap of missing values, user == {user}")
No description has been provided for this image

Based on the above plots it seems like there is some missingness that needs to be imputed

Take an example user and impute the missingness in their data

In [31]:
user = top_users[0]
user
Out[31]:
'Moti156'
In [77]:
plot_df = df.select_dtypes(np.number).loc[(user, ), :]
# plot_df.index.min(), plot_df.index.max()
new_index = pd.date_range(start=plot_df.index.min(), end=plot_df.index.max())
plot_df = plot_df.reindex(new_index)
# plot_df.index = plot_df.index.astype(str)
In [78]:
feature_selection = plot_df.columns.tolist()
In [79]:
current_feature = "absorption"
# temp = plot_df.isna().astype(int)
temp = plot_df
x = temp[current_feature]
In [80]:
_ = plt.figure(figsize=(20, 5))
_ = plt.plot(x, label=current_feature)
_ = plt.scatter(x.index, x.values, c="green")
_ = plt.xticks(rotation=90)
_ = plt.legend()
No description has been provided for this image
In [81]:
impute_estimator = ExtraTreesRegressor(n_estimators=2, random_state=0)
estimator = make_pipeline(
        IterativeImputer(random_state=0, estimator=impute_estimator)
    )

ts_imp_df = pd.DataFrame(estimator.fit_transform(temp.loc[:, feature_selection].values), 
                           index = temp.index, 
                           columns = feature_selection)
/opt/miniconda3/envs/general/lib/python3.9/site-packages/sklearn/impute/_iterative.py:685: ConvergenceWarning: [IterativeImputer] Early stopping criterion not reached.
  warnings.warn("[IterativeImputer] Early stopping criterion not"
In [84]:
ts_imp_df = temp.loc[:, feature_selection].apply(lambda x: x.interpolate(method='polynomial', order=5))
In [85]:
_ = plt.figure(figsize=(20, 5))
_ = plt.plot(ts_imp_df[current_feature], label=current_feature)
_ = plt.scatter(ts_imp_df[current_feature].index, ts_imp_df[current_feature].values, c="red")
_ = plt.scatter(x.index, x.values, c="green")
_ = plt.xticks(rotation=90)
_ = plt.legend()
No description has been provided for this image
In [72]:
cmap_name = "neurocast_colours"
colors = [ 
    "#FFFFFF",
          "#2768c1"
          ]
n_bin = 10
cm = LinearSegmentedColormap.from_list(cmap_name, colors, N=n_bin)
In [73]:
x = ts_imp_df[current_feature]
In [74]:
example_rm = RecurrencePlot(time_series=x.values, metric="manhattan", dim=3, tau=2, recurrence_rate=0.05).recurrence_matrix()

# rotate the matrix 90 degrees so it starts at day 0 in the bottom left corner
new_matrix = [[example_rm[j][i] for j in range(len(example_rm))] for i in range(len(example_rm[0])-1,-1,-1)]

fig, axes = plt.subplots(2, 1, figsize=(7, 9), gridspec_kw={'height_ratios': [1, 3]}) 
sns.despine(left=False)

axe = sns.lineplot(x=np.arange(0, ts_imp_df.loc[:, current_feature].shape[0]), 
                 y=ts_imp_df.loc[:, current_feature], 
                color = "#2768c1",
                ax=axes[0])
# axe.set_ylabel(None)
# axe.set_ylabel(r"$\tilde{u}\left(HT_{n}\right)$")

# axe.set_ylabel(r'$\tilde{u}\left(HT_{n}\right) \, \, \left[\mathrm{ms} \right]$')
axe.set_title(f"Recurrence Plot User == {user}")


yticks = list(np.linspace(0, len(x)-3, 20, dtype=np.int))
_xticklabels = list(np.arange(0, len(x), 1))
xticklabels = [_xticklabels[idx] for idx in yticks]
_yticklabels = list(np.arange(len(x)-3, -1, -1))
yticklabels = [_yticklabels[idx] for idx in yticks]

ax = sns.heatmap(new_matrix, cmap=cm, ax=axes[1], cbar=False, xticklabels = xticklabels, yticklabels = yticklabels
                )
_ = ax.set_xticks(yticks)
_ = ax.set_yticks(yticks)
Calculating recurrence plot at fixed recurrence rate...
Calculating the manhattan distance matrix...
No description has been provided for this image

To summarise the above plots it is unlikely we will have very trustworthy results if we would continue to use NLTSA methods on this data set for the following reasons:

  • The time series are not really long enough (max row amount == 64 for user == Moti148)
  • The time series are not continuous/ are patchy with missingness. For example for user Moti148 the row amount was 64 over a date range of 128 days
  • Attempting to impute these missing values is possible, however looking at the output (time series plot of absorption for user == Moti156), I am dubious that this is a good idea