424 KiB
424 KiB
Attempting a permutation approach based on matching demographics in users and comparing the mask based detertminants in the citizen shield project¶
Discussion¶
- Main point is that the evolutionary potential of a system depends on its current state
- I'm considering applications to psychology/behaviour
- Mainly: a dataset has people as rows and features as columns, Likert scale. Let's say the columns are three behaviours and five demographics.
- For every row, I want to find the most similar row in terms of demographics and the combination of behaviours.
- Let's say more behaviours are better, so I want to find a slightly superior comparison person to everyone, especially for those low in the behaviours
- So in theory, the person could be nudged to the superior 3-dimensional behavioural state, even though their 5-dimensional demographics are unchangable
- (Three and five coming from here)
- ... and least energy cost would be a switch to the most adjacent state
- A new variable would be, for every person, the distance to their closest adjacent superior
- Actually, you could just group the data by permutations of the 5 demographics, for each take a distance matrix of the 3 behaviours, and -- for each row -- pull the adjacent row numbers which are not identical
- Is this a stupid idea?
- Could then look at distance distributions of the groups (i.e. permutations of the 5 demographics), and a uniform distribution would be the nicest -- everyone in that group would have an improved version of themselves, so we "know" based on the data, that it's not impossible for someone with those demographics to improve behaviours
- A bimodal distribution would be problematic, as there'd be a potentially uncrossable chasm for the people in the lower behavioural mode
- So interventions could be focused on those groups, where individuals have potential for large stepwise improvement
In [1]:
import os
In [22]:
import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sns from jmspack.utils import JmsColors, flatten from itertools import product from scipy.spatial.distance import pdist, squareform from pyxdameraulevenshtein import damerau_levenshtein_distance as DLD
In [3]:
if "jms_style_sheet" in plt.style.available: _ = plt.style.use("jms_style_sheet")
In [4]:
df = (pd.read_csv("kp_determinants2/kp_determinants_and_demographics.csv") .dropna() )
In [5]:
df.info()
<class 'pandas.core.frame.DataFrame'> Int64Index: 2407 entries, 0 to 2458 Data columns (total 31 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 demographic_gender 2407 non-null int64 1 demographic_age 2407 non-null int64 2 demographic_region 2407 non-null float64 3 demographic_education 2407 non-null float64 4 demographic_living_with 2407 non-null float64 5 demographic_underage_children 2407 non-null object 6 demographic_income 2407 non-null float64 7 effectiveness_masks 2407 non-null int64 8 effectiveness_hometests 2407 non-null int64 9 effectiveness_ventilation 2407 non-null int64 10 effectiveness_quarantine 2407 non-null int64 11 effectiveness_avoid_meeting 2407 non-null int64 12 effectiveness_announce_test_result 2407 non-null int64 13 effectiveness_avoid_public_events 2407 non-null int64 14 effectiveness_covid_passport 2407 non-null int64 15 ease_masks 2407 non-null int64 16 ease_hometests 2407 non-null int64 17 ease_ventilation 2407 non-null int64 18 ease_quarantine 2407 non-null int64 19 ease_avoid_meeting 2407 non-null int64 20 ease_announce_test_result 2407 non-null int64 21 ease_avoid_public_events 2407 non-null int64 22 ease_covid_passport 2407 non-null int64 23 persistence_masks 2407 non-null int64 24 persistence_hometests 2407 non-null int64 25 persistence_ventilation 2407 non-null int64 26 persistence_quarantine 2407 non-null int64 27 persistence_avoid_meeting 2407 non-null int64 28 persistence_announce_test_result 2407 non-null int64 29 persistence_avoid_public_events 2407 non-null int64 30 persistence_covid_passport 2407 non-null int64 dtypes: float64(4), int64(26), object(1) memory usage: 601.8+ KB
In [6]:
display(df.head()); print(df.shape)
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| demographic_gender | demographic_age | demographic_region | demographic_education | demographic_living_with | demographic_underage_children | demographic_income | effectiveness_masks | effectiveness_hometests | effectiveness_ventilation | ... | ease_avoid_public_events | ease_covid_passport | persistence_masks | persistence_hometests | persistence_ventilation | persistence_quarantine | persistence_avoid_meeting | persistence_announce_test_result | persistence_avoid_public_events | persistence_covid_passport | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1 | 7 | 17.0 | 2.0 | 2.0 | Yes | 3.0 | 2 | 1 | 5 | ... | 5 | 0 | 3 | 4 | 5 | 4 | 4 | 5 | 5 | 0 |
| 1 | 1 | 4 | 17.0 | 3.0 | 2.0 | Yes | 2.0 | 3 | 1 | 5 | ... | 5 | 5 | 2 | 1 | 5 | 1 | 1 | 5 | 4 | 5 |
| 2 | 1 | 6 | 17.0 | 4.0 | 2.0 | Yes | 3.0 | 5 | 5 | 5 | ... | 2 | 4 | 3 | 4 | 5 | 3 | 3 | 5 | 1 | 4 |
| 3 | 1 | 7 | 17.0 | 4.0 | 2.0 | Yes | 3.0 | 5 | 5 | 4 | ... | 2 | 5 | 5 | 5 | 5 | 5 | 3 | 5 | 3 | 5 |
| 4 | 1 | 4 | 17.0 | 2.0 | 1.0 | No | 4.0 | 1 | 1 | 3 | ... | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
5 rows × 31 columns
(2407, 31)
In [7]:
demo_columns = df.filter(regex="demographic").drop(["demographic_region", "demographic_age"], axis=1).columns.tolist() # demo_columns = df.filter(regex="demographic").drop(["demographic_region"], axis=1).columns.tolist() # demo_columns = ["demographic_gender", "demographic_education"] # feature_columns = df.filter(regex="ease").columns.tolist() feature_columns = df.filter(regex="effectiveness").columns.tolist() # feature_columns = df.filter(regex="persistence").columns.tolist()
Create a demo_options_df¶
In [8]:
df[demo_columns].nunique()
Out[8]:
demographic_gender 3 demographic_education 5 demographic_living_with 3 demographic_underage_children 2 demographic_income 5 dtype: int64
In [9]:
options_list = [list(x) for x in df[demo_columns].apply(lambda col: col.unique()).tolist()]
In [10]:
all_options_list = list(product(*options_list))
In [11]:
options_df = (pd.DataFrame(all_options_list, columns=demo_columns) .reset_index() .rename(columns={"index": "option_number"})) options_df.head()
Out[11]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| option_number | demographic_gender | demographic_education | demographic_living_with | demographic_underage_children | demographic_income | |
|---|---|---|---|---|---|---|
| 0 | 0 | 1 | 2.0 | 2.0 | Yes | 3.0 |
| 1 | 1 | 1 | 2.0 | 2.0 | Yes | 2.0 |
| 2 | 2 | 1 | 2.0 | 2.0 | Yes | 4.0 |
| 3 | 3 | 1 | 2.0 | 2.0 | Yes | 1.0 |
| 4 | 4 | 1 | 2.0 | 2.0 | Yes | 5.0 |
In [12]:
df_perm = pd.DataFrame() for row in options_df.index: tmp = pd.merge(options_df.iloc[[row], :], df[demo_columns + feature_columns], on=demo_columns) df_perm = pd.concat([df_perm, tmp])
In [13]:
amount_cutoff = 20 amount_top_demos = 10 plot_df = (df_perm["option_number"] .value_counts() .to_frame() .reset_index() .rename(columns={"index": "option_number", "option_number": "frequency"}) .pipe(lambda d: d[d["frequency"]>=amount_cutoff]) ) top_demos_list = plot_df["option_number"].head(amount_top_demos).values.tolist() display(plot_df.head(2)); print(plot_df.shape)
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| option_number | frequency | |
|---|---|---|
| 0 | 155 | 142 |
| 1 | 185 | 97 |
(40, 2)
In [14]:
_ = plt.figure(figsize=(2, 8)) _ = sns.heatmap(data=plot_df.set_index("option_number"), cmap="Blues", annot=True, fmt=".0f", cbar=False)
In [15]:
plot_df = (df_perm.pipe(lambda d: d[d["option_number"].isin([155, 185])]) .assign(**{"option_number": lambda d: d["option_number"].astype("str")})) plot_df.head(2)
Out[15]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| option_number | demographic_gender | demographic_education | demographic_living_with | demographic_underage_children | demographic_income | effectiveness_masks | effectiveness_hometests | effectiveness_ventilation | effectiveness_quarantine | effectiveness_avoid_meeting | effectiveness_announce_test_result | effectiveness_avoid_public_events | effectiveness_covid_passport | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 155 | 2 | 2.0 | 2.0 | No | 3.0 | 5 | 0 | 5 | 5 | 4 | 5 | 5 | 5 |
| 1 | 155 | 2 | 2.0 | 2.0 | No | 3.0 | 5 | 5 | 5 | 5 | 4 | 5 | 4 | 5 |
In [16]:
_ = plt.figure(figsize=(30, 5)) _ = sns.swarmplot(data=plot_df, x=feature_columns[2], y=feature_columns[1], hue="option_number")
In [17]:
selection_df = df_perm.pipe(lambda d: d[d["option_number"].isin(top_demos_list)]) plot_df = (selection_df .drop(demo_columns, axis=1) .sort_values(by=["option_number"] + feature_columns) .groupby("option_number") .diff(axis=0) .mean(axis=1) .to_frame("feature_diff_mean") .assign(**{"option_number": selection_df["option_number"].values}) .dropna() .assign(**{"option_number": lambda d: d["option_number"].astype("category")}) ) plot_df.head()
Out[17]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| feature_diff_mean | option_number | |
|---|---|---|
| 68 | 0.750 | 5 |
| 88 | 0.250 | 5 |
| 78 | 0.250 | 5 |
| 2 | -0.250 | 5 |
| 31 | 0.625 | 5 |
In [18]:
if amount_top_demos <= 10: show_legend=True else: show_legend=False
In [19]:
_ = plt.figure(figsize=(15, 5)) _ = sns.histplot(data=plot_df, x="feature_diff_mean", hue="option_number", kde=True, bins=50, legend=show_legend)
In [20]:
plot_df = (selection_df .drop(demo_columns, axis=1) .sort_values(by=["option_number"] + feature_columns) .groupby("option_number") .diff(axis=0) .assign(**{"option_number": selection_df["option_number"].values}) .dropna() .assign(**{"option_number": lambda d: d["option_number"].astype("category")}) .melt(id_vars="option_number") ) plot_df.head()
Out[20]:
<style scoped="">
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| option_number | variable | value | |
|---|---|---|---|
| 0 | 5 | effectiveness_masks | 1.0 |
| 1 | 5 | effectiveness_masks | 0.0 |
| 2 | 5 | effectiveness_masks | 1.0 |
| 3 | 5 | effectiveness_masks | 0.0 |
| 4 | 5 | effectiveness_masks | 0.0 |
In [21]:
_ = plt.figure(figsize=(15, 5)) _ = sns.stripplot(data=plot_df, x="variable", y="value", hue="option_number", dodge=True, legend=show_legend) _ = plt.xticks(rotation=90)
In [58]:
DLD_df = pd.DataFrame() for option in top_demos_list: tmp_df = (selection_df .pipe(lambda d: d[d["option_number"]==option]) .assign(**{"combined_score": lambda d: d[feature_columns] .astype(str) .sum(axis=1) .astype(int)}) .loc[:, ["option_number", "combined_score"]] .sort_values(by=["option_number", "combined_score"]) .assign(**{"combined_score": lambda d: d["combined_score"].astype(str)}) .reset_index(drop=True) ) tmp_df["DLD"] = [np.nan] + [DLD(tmp_df.loc[x, "combined_score"], tmp_df.loc[x+1, "combined_score"]) for x in tmp_df.index.tolist()[:-1]] DLD_df = pd.concat([DLD_df, tmp_df])
In [63]:
_ = plt.figure(figsize=(5, 5)) _ = sns.histplot(data=DLD_df.dropna().assign(**{"option_number": lambda d: d["option_number"].astype(str)}), x="DLD", hue="option_number", kde=True, bins=50, legend=show_legend)