2.3 MiB
2.3 MiB
Work Motivation Under Lockdown Exploratory Data Analysis (EDA)¶
Two datasets...
Post-workday evaluations of work motivation related stuff. Large spread across how many time points people collected; few decent-ish ones.
Traditional questionnaires from the same participants.
- Allocation indicates whether they took part in our motivation self-management training during the period or belonged to a waitlist control group.
- Pre-, mid- and post-questionnaires --> TIME == 1, 2 and 3 respectively.
- Everything from intrinsic1 to Psycap5 are raw answers. From Intrinsic to Optimism are sumscores.
- This was a feasibility & acceptability study so it's not powered for or aiming to look at outcomes. But if there was one, it could be the "Relative Autonomy Index" (RAI), a sumscore conventionally built as:
- Amotivation * -3 + ExternalTotal * -2 + Introjected * -1 + Identified * 2 + Intrinsic * 3
- If I remember correctly, there were some people whose RAI was boosted a lot
- In general:
- Good things to go up during the intervention (or life in general):
- Intrinsic, Identified
- Competence, Relatedness, Autonomy
- Bad things to go up during the intervention (or life in general):
- Amotivation, External (material & social), Introjected
- AutonomyThwarting, RelatednessThwarting, CompetenceThwarting
- Unsure if these were named in the data, or just some of the items named autonomyx/relatednessx/competencex
- Good things to go up during the intervention (or life in general):
This EDA only addresses the daily data, specifically data=="data/moti_feasibility_james_daily.csv"¶
In [1]:
import pandas as pd import numpy as np import seaborn as sns import matplotlib.pyplot as plt from matplotlib.colors import LinearSegmentedColormap import session_info from jmspack.NLTSA import (flatten, ts_levels, fluctuation_intensity, distribution_uniformity, complexity_resonance, complexity_resonance_diagram) from sklearn.preprocessing import MinMaxScaler # To use this experimental feature, we need to explicitly ask for it: from sklearn.experimental import enable_iterative_imputer # noqa from sklearn.impute import IterativeImputer from sklearn.tree import DecisionTreeRegressor from sklearn.ensemble import ExtraTreesRegressor from sklearn.neighbors import KNeighborsRegressor from sklearn.pipeline import make_pipeline from pyunicorn.timeseries import RecurrencePlot
pyunicorn: Package netCDF4 could not be loaded. Some functionality in class Data might not be available! pyunicorn: Package netCDF4 could not be loaded. Some functionality in class NetCDFDictionary might not be available!
Show the session information of the packages used in this analysis
In [2]:
session_info.show(write_req_file=False, req_file_name="work_motivation_EDA_requirements.txt",)
Out[2]:
Click to view session information
----- jmspack 0.0.3 matplotlib 3.3.4 numpy 1.19.2 pandas 1.2.3 pyunicorn NA seaborn 0.11.1 session_info 1.0.0 sklearn 0.24.1 -----
Click to view modules imported as dependencies
PIL 8.1.2 appnope 0.1.2 backcall 0.2.0 cffi 1.14.5 colorama 0.4.4 cycler 0.10.0 cython_runtime NA dateutil 2.8.1 decorator 4.4.2 igraph 0.9.1 ipykernel 5.3.4 ipython_genutils 0.2.0 ipywidgets 7.6.3 jedi 0.17.2 joblib 0.17.0 kiwisolver 1.3.1 mpl_toolkits NA parso 0.7.0 pexpect 4.8.0 pickleshare 0.7.5 pkg_resources NA prompt_toolkit 3.0.8 ptyprocess 0.7.0 pyexpat NA pygments 2.8.1 pyparsing 2.4.7 pytz 2021.1 scipy 1.6.1 six 1.15.0 statsmodels 0.12.2 storemagic NA texttable 1.6.3 tornado 6.1 traitlets 5.0.5 wcwidth 0.2.5 zmq 20.0.0
----- IPython 7.21.0 jupyter_client 6.1.7 jupyter_core 4.7.1 jupyterlab 2.2.6 notebook 6.2.0 ----- Python 3.9.2 (default, Mar 3 2021, 11:58:52) [Clang 10.0.0 ] macOS-10.16-x86_64-i386-64bit ----- Session information updated at 2021-06-10 12:27
Read in the raw daily data¶
In [3]:
_df = pd.read_csv("data/moti_feasibility_james_daily.csv")
In [4]:
_df = _df.assign(date=pd.to_datetime(_df["Date (submitted)"].str.split("T", expand=True)[0], format="%Y-%m-%dT%H:%M:%S"))
In [5]:
def convert_numeric_where_possible(x): try: return x.astype(np.number) except: return x
Reshape the data frame and convert data types to numeric where possible¶
In [6]:
df = (_df .loc[:, ["User", "Field", "Value", "date"]] .pivot(index=["User", "date"], columns="Field") .droplevel(level=0, axis=1) .apply(convert_numeric_where_possible) ) df.head(2)
Out[6]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| Field | DailySituation | absorption | amotivation | autonomy | competence | dedication | emotional_drain | enjoyment | external pressure | howami_participation_ended | howami_participation_started | importance | interest | internal pressure | productivity_work | relatedness | satisfaction_work | strategies | time_pressure | vigor | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| User | date | ||||||||||||||||||||
| Moti105 | 2020-03-03 | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | 2020-03-03 | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN | NaN |
| 2020-03-09 | NaN | 68.0 | 66.0 | 33.0 | 62.0 | 69.0 | NaN | 31.0 | 52.0 | NaN | NaN | 67.0 | 74.0 | 60.0 | 72.0 | 65.0 | 42.0 | emphasize_autonomy;reflect_desire;identifyas_r... | NaN | 64.0 |
In [7]:
df.info()
<class 'pandas.core.frame.DataFrame'>
MultiIndex: 1674 entries, ('Moti105', Timestamp('2020-03-03 00:00:00')) to ('Moti218', Timestamp('2020-05-15 00:00:00'))
Data columns (total 20 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 DailySituation 1173 non-null object
1 absorption 1487 non-null float64
2 amotivation 1487 non-null float64
3 autonomy 1487 non-null float64
4 competence 1487 non-null float64
5 dedication 1487 non-null float64
6 emotional_drain 1126 non-null float64
7 enjoyment 1487 non-null float64
8 external pressure 1487 non-null float64
9 howami_participation_ended 11 non-null object
10 howami_participation_started 101 non-null object
11 importance 1487 non-null float64
12 interest 1487 non-null float64
13 internal pressure 1487 non-null float64
14 productivity_work 1487 non-null float64
15 relatedness 1487 non-null float64
16 satisfaction_work 1487 non-null float64
17 strategies 1301 non-null object
18 time_pressure 1126 non-null float64
19 vigor 1487 non-null float64
dtypes: float64(16), object(4)
memory usage: 274.8+ KB
Count the amount of rows per user and plot¶
The aim of this is to assess whether there is enough data to do some of the more complext NLTSA methods
In [8]:
row_amount_df = (df .reset_index() .groupby("User") .count() .loc[:, ["date"]] .rename(columns={"date": "row_amount"}) .sort_values(by="row_amount") .reset_index())
In [9]:
row_amount_df.tail(1)
Out[9]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| Field | User | row_amount |
|---|---|---|
| 101 | Moti148 | 64 |
In [10]:
_ = plt.figure(figsize=(20, 4)) _ = sns.barplot(data=row_amount_df, x="User", y="row_amount") _ = plt.xticks(rotation=90) _ = plt.axhline(30, c="red", ls="--", label="30 day cutoff") _ = plt.title("Row amounts per user") _ = plt.legend()
Plot the heatmaps of the top 5 users with the highest amount of rows¶
The aim of this is to assess the amount of missingness time wise (i.e. the user could have a lot of data, spread over a long period with large gaps in the middle).
In [11]:
top_users = row_amount_df.tail(5).User.tolist()
In [25]:
scale_data = True for user in top_users: plot_df = df.select_dtypes(np.number).loc[(user, ), :] # plot_df.index.min(), plot_df.index.max() new_index = pd.date_range(start=plot_df.index.min(), end=plot_df.index.max()) plot_df = plot_df.reindex(new_index) if scale_data: plot_df = pd.DataFrame(MinMaxScaler().fit_transform(plot_df), index=plot_df.index, columns=plot_df.columns) plot_df.index = plot_df.index.astype(str) _ = plt.figure(figsize=(20, 5)) _ = sns.lineplot(data = plot_df.reset_index().melt(id_vars="index"), x="index", y="value", hue="Field") _ = sns.scatterplot(data = plot_df.reset_index().melt(id_vars="index"), x="index", y="value", hue="Field", legend=False) _ = plt.xticks(rotation=90) _ = plt.title(f"Lineplot of raw values, user == {user}") _ = plt.figure(figsize=(20, 5)) _ = sns.heatmap(plot_df.T) _ = plt.title(f"Heatmap of raw values, user == {user}")
Show the heatmap of the missing values for the last user plotted above¶
In [26]:
_ = plt.figure(figsize=(20, 5)) _ = sns.heatmap(plot_df.isna().astype(int).T) _ = plt.title(f"Heatmap of missing values, user == {user}")
Based on the above plots it seems like there is some missingness that needs to be imputed¶
Take an example user and impute the missingness in their data
In [31]:
user = top_users[0] user
Out[31]:
'Moti156'
In [77]:
plot_df = df.select_dtypes(np.number).loc[(user, ), :] # plot_df.index.min(), plot_df.index.max() new_index = pd.date_range(start=plot_df.index.min(), end=plot_df.index.max()) plot_df = plot_df.reindex(new_index) # plot_df.index = plot_df.index.astype(str)
In [78]:
feature_selection = plot_df.columns.tolist()
In [79]:
current_feature = "absorption" # temp = plot_df.isna().astype(int) temp = plot_df x = temp[current_feature]
In [80]:
_ = plt.figure(figsize=(20, 5)) _ = plt.plot(x, label=current_feature) _ = plt.scatter(x.index, x.values, c="green") _ = plt.xticks(rotation=90) _ = plt.legend()
In [81]:
impute_estimator = ExtraTreesRegressor(n_estimators=2, random_state=0) estimator = make_pipeline( IterativeImputer(random_state=0, estimator=impute_estimator) ) ts_imp_df = pd.DataFrame(estimator.fit_transform(temp.loc[:, feature_selection].values), index = temp.index, columns = feature_selection)
/opt/miniconda3/envs/general/lib/python3.9/site-packages/sklearn/impute/_iterative.py:685: ConvergenceWarning: [IterativeImputer] Early stopping criterion not reached.
warnings.warn("[IterativeImputer] Early stopping criterion not"
In [84]:
ts_imp_df = temp.loc[:, feature_selection].apply(lambda x: x.interpolate(method='polynomial', order=5))
In [85]:
_ = plt.figure(figsize=(20, 5)) _ = plt.plot(ts_imp_df[current_feature], label=current_feature) _ = plt.scatter(ts_imp_df[current_feature].index, ts_imp_df[current_feature].values, c="red") _ = plt.scatter(x.index, x.values, c="green") _ = plt.xticks(rotation=90) _ = plt.legend()
In [72]:
cmap_name = "neurocast_colours" colors = [ "#FFFFFF", "#2768c1" ] n_bin = 10 cm = LinearSegmentedColormap.from_list(cmap_name, colors, N=n_bin)
In [73]:
x = ts_imp_df[current_feature]
In [74]:
example_rm = RecurrencePlot(time_series=x.values, metric="manhattan", dim=3, tau=2, recurrence_rate=0.05).recurrence_matrix() # rotate the matrix 90 degrees so it starts at day 0 in the bottom left corner new_matrix = [[example_rm[j][i] for j in range(len(example_rm))] for i in range(len(example_rm[0])-1,-1,-1)] fig, axes = plt.subplots(2, 1, figsize=(7, 9), gridspec_kw={'height_ratios': [1, 3]}) sns.despine(left=False) axe = sns.lineplot(x=np.arange(0, ts_imp_df.loc[:, current_feature].shape[0]), y=ts_imp_df.loc[:, current_feature], color = "#2768c1", ax=axes[0]) # axe.set_ylabel(None) # axe.set_ylabel(r"$\tilde{u}\left(HT_{n}\right)$") # axe.set_ylabel(r'$\tilde{u}\left(HT_{n}\right) \, \, \left[\mathrm{ms} \right]$') axe.set_title(f"Recurrence Plot User == {user}") yticks = list(np.linspace(0, len(x)-3, 20, dtype=np.int)) _xticklabels = list(np.arange(0, len(x), 1)) xticklabels = [_xticklabels[idx] for idx in yticks] _yticklabels = list(np.arange(len(x)-3, -1, -1)) yticklabels = [_yticklabels[idx] for idx in yticks] ax = sns.heatmap(new_matrix, cmap=cm, ax=axes[1], cbar=False, xticklabels = xticklabels, yticklabels = yticklabels ) _ = ax.set_xticks(yticks) _ = ax.set_yticks(yticks)
Calculating recurrence plot at fixed recurrence rate... Calculating the manhattan distance matrix...
To summarise the above plots it is unlikely we will have very trustworthy results if we would continue to use NLTSA methods on this data set for the following reasons:¶
- The time series are not really long enough (max row amount == 64 for user == Moti148)
- The time series are not continuous/ are patchy with missingness. For example for user Moti148 the row amount was 64 over a date range of 128 days
- Attempting to impute these missing values is possible, however looking at the output (time series plot of absorption for user == Moti156), I am dubious that this is a good idea