Files
matti_jms_collabs/citizen_shield/citizen_shield_permutation_approach.ipynb
T

424 KiB
Raw Blame History

Attempting a permutation approach based on matching demographics in users and comparing the mask based detertminants in the citizen shield project

Discussion

  • Main point is that the evolutionary potential of a system depends on its current state
  • I'm considering applications to psychology/behaviour
  • Mainly: a dataset has people as rows and features as columns, Likert scale. Let's say the columns are three behaviours and five demographics.
  • For every row, I want to find the most similar row in terms of demographics and the combination of behaviours.
  • Let's say more behaviours are better, so I want to find a slightly superior comparison person to everyone, especially for those low in the behaviours
  • So in theory, the person could be nudged to the superior 3-dimensional behavioural state, even though their 5-dimensional demographics are unchangable
  • (Three and five coming from here)
  • ... and least energy cost would be a switch to the most adjacent state
  • A new variable would be, for every person, the distance to their closest adjacent superior
  • Actually, you could just group the data by permutations of the 5 demographics, for each take a distance matrix of the 3 behaviours, and -- for each row -- pull the adjacent row numbers which are not identical
  • Is this a stupid idea?
  • Could then look at distance distributions of the groups (i.e. permutations of the 5 demographics), and a uniform distribution would be the nicest -- everyone in that group would have an improved version of themselves, so we "know" based on the data, that it's not impossible for someone with those demographics to improve behaviours
  • A bimodal distribution would be problematic, as there'd be a potentially uncrossable chasm for the people in the lower behavioural mode
  • So interventions could be focused on those groups, where individuals have potential for large stepwise improvement
In [1]:
import os
In [22]:
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from jmspack.utils import JmsColors, flatten
from itertools import product
from scipy.spatial.distance import pdist, squareform
from pyxdameraulevenshtein import damerau_levenshtein_distance as DLD
In [3]:
if "jms_style_sheet" in plt.style.available:
    _ = plt.style.use("jms_style_sheet")
In [4]:
df = (pd.read_csv("kp_determinants2/kp_determinants_and_demographics.csv")
      .dropna()
     )
In [5]:
df.info()
<class 'pandas.core.frame.DataFrame'>
Int64Index: 2407 entries, 0 to 2458
Data columns (total 31 columns):
 #   Column                              Non-Null Count  Dtype  
---  ------                              --------------  -----  
 0   demographic_gender                  2407 non-null   int64  
 1   demographic_age                     2407 non-null   int64  
 2   demographic_region                  2407 non-null   float64
 3   demographic_education               2407 non-null   float64
 4   demographic_living_with             2407 non-null   float64
 5   demographic_underage_children       2407 non-null   object 
 6   demographic_income                  2407 non-null   float64
 7   effectiveness_masks                 2407 non-null   int64  
 8   effectiveness_hometests             2407 non-null   int64  
 9   effectiveness_ventilation           2407 non-null   int64  
 10  effectiveness_quarantine            2407 non-null   int64  
 11  effectiveness_avoid_meeting         2407 non-null   int64  
 12  effectiveness_announce_test_result  2407 non-null   int64  
 13  effectiveness_avoid_public_events   2407 non-null   int64  
 14  effectiveness_covid_passport        2407 non-null   int64  
 15  ease_masks                          2407 non-null   int64  
 16  ease_hometests                      2407 non-null   int64  
 17  ease_ventilation                    2407 non-null   int64  
 18  ease_quarantine                     2407 non-null   int64  
 19  ease_avoid_meeting                  2407 non-null   int64  
 20  ease_announce_test_result           2407 non-null   int64  
 21  ease_avoid_public_events            2407 non-null   int64  
 22  ease_covid_passport                 2407 non-null   int64  
 23  persistence_masks                   2407 non-null   int64  
 24  persistence_hometests               2407 non-null   int64  
 25  persistence_ventilation             2407 non-null   int64  
 26  persistence_quarantine              2407 non-null   int64  
 27  persistence_avoid_meeting           2407 non-null   int64  
 28  persistence_announce_test_result    2407 non-null   int64  
 29  persistence_avoid_public_events     2407 non-null   int64  
 30  persistence_covid_passport          2407 non-null   int64  
dtypes: float64(4), int64(26), object(1)
memory usage: 601.8+ KB
In [6]:
display(df.head()); print(df.shape)
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
demographic_gender demographic_age demographic_region demographic_education demographic_living_with demographic_underage_children demographic_income effectiveness_masks effectiveness_hometests effectiveness_ventilation ... ease_avoid_public_events ease_covid_passport persistence_masks persistence_hometests persistence_ventilation persistence_quarantine persistence_avoid_meeting persistence_announce_test_result persistence_avoid_public_events persistence_covid_passport
0 1 7 17.0 2.0 2.0 Yes 3.0 2 1 5 ... 5 0 3 4 5 4 4 5 5 0
1 1 4 17.0 3.0 2.0 Yes 2.0 3 1 5 ... 5 5 2 1 5 1 1 5 4 5
2 1 6 17.0 4.0 2.0 Yes 3.0 5 5 5 ... 2 4 3 4 5 3 3 5 1 4
3 1 7 17.0 4.0 2.0 Yes 3.0 5 5 4 ... 2 5 5 5 5 5 3 5 3 5
4 1 4 17.0 2.0 1.0 No 4.0 1 1 3 ... 1 1 1 1 1 1 1 1 1 1

5 rows × 31 columns

(2407, 31)
In [7]:
demo_columns = df.filter(regex="demographic").drop(["demographic_region", "demographic_age"], axis=1).columns.tolist()
# demo_columns = df.filter(regex="demographic").drop(["demographic_region"], axis=1).columns.tolist()
# demo_columns = ["demographic_gender", "demographic_education"]
# feature_columns = df.filter(regex="ease").columns.tolist()
feature_columns = df.filter(regex="effectiveness").columns.tolist()
# feature_columns = df.filter(regex="persistence").columns.tolist()

Create a demo_options_df

In [8]:
df[demo_columns].nunique()
Out[8]:
demographic_gender               3
demographic_education            5
demographic_living_with          3
demographic_underage_children    2
demographic_income               5
dtype: int64
In [9]:
options_list = [list(x) for x in df[demo_columns].apply(lambda col: col.unique()).tolist()]
In [10]:
all_options_list = list(product(*options_list))
In [11]:
options_df = (pd.DataFrame(all_options_list, columns=demo_columns)
              .reset_index()
              .rename(columns={"index": "option_number"}))
options_df.head()
Out[11]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
option_number demographic_gender demographic_education demographic_living_with demographic_underage_children demographic_income
0 0 1 2.0 2.0 Yes 3.0
1 1 1 2.0 2.0 Yes 2.0
2 2 1 2.0 2.0 Yes 4.0
3 3 1 2.0 2.0 Yes 1.0
4 4 1 2.0 2.0 Yes 5.0
In [12]:
df_perm = pd.DataFrame()
for row in options_df.index:
    tmp = pd.merge(options_df.iloc[[row], :], df[demo_columns + feature_columns], on=demo_columns)
    df_perm = pd.concat([df_perm, tmp])
In [13]:
amount_cutoff = 20
amount_top_demos = 10
plot_df = (df_perm["option_number"]
.value_counts()
.to_frame()
.reset_index()
.rename(columns={"index": "option_number", "option_number": "frequency"})
.pipe(lambda d: d[d["frequency"]>=amount_cutoff])
)
top_demos_list = plot_df["option_number"].head(amount_top_demos).values.tolist()
display(plot_df.head(2)); print(plot_df.shape)
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
option_number frequency
0 155 142
1 185 97
(40, 2)
In [14]:
_ = plt.figure(figsize=(2, 8))
_ = sns.heatmap(data=plot_df.set_index("option_number"), cmap="Blues", annot=True, fmt=".0f", cbar=False)
No description has been provided for this image
In [15]:
plot_df = (df_perm.pipe(lambda d: d[d["option_number"].isin([155, 185])])
           .assign(**{"option_number": lambda d: d["option_number"].astype("str")}))
plot_df.head(2)
Out[15]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
option_number demographic_gender demographic_education demographic_living_with demographic_underage_children demographic_income effectiveness_masks effectiveness_hometests effectiveness_ventilation effectiveness_quarantine effectiveness_avoid_meeting effectiveness_announce_test_result effectiveness_avoid_public_events effectiveness_covid_passport
0 155 2 2.0 2.0 No 3.0 5 0 5 5 4 5 5 5
1 155 2 2.0 2.0 No 3.0 5 5 5 5 4 5 4 5
In [16]:
_ = plt.figure(figsize=(30, 5))
_ = sns.swarmplot(data=plot_df, x=feature_columns[2], y=feature_columns[1], hue="option_number")
No description has been provided for this image
In [17]:
selection_df = df_perm.pipe(lambda d: d[d["option_number"].isin(top_demos_list)])
plot_df = (selection_df
 .drop(demo_columns, axis=1)
 .sort_values(by=["option_number"] + feature_columns)
 .groupby("option_number")
 .diff(axis=0)
 .mean(axis=1)
 .to_frame("feature_diff_mean")
 .assign(**{"option_number": selection_df["option_number"].values})
  .dropna()
 .assign(**{"option_number": lambda d: d["option_number"].astype("category")})
 )
plot_df.head()
Out[17]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
feature_diff_mean option_number
68 0.750 5
88 0.250 5
78 0.250 5
2 -0.250 5
31 0.625 5
In [18]:
if amount_top_demos <= 10:
    show_legend=True
else:
    show_legend=False
In [19]:
_ = plt.figure(figsize=(15, 5))
_ = sns.histplot(data=plot_df, 
                 x="feature_diff_mean", 
                 hue="option_number", 
                 kde=True, bins=50, 
                 legend=show_legend)
No description has been provided for this image
In [20]:
plot_df = (selection_df
 .drop(demo_columns, axis=1)
 .sort_values(by=["option_number"] + feature_columns)
 .groupby("option_number")
 .diff(axis=0)
 .assign(**{"option_number": selection_df["option_number"].values})
  .dropna()
 .assign(**{"option_number": lambda d: d["option_number"].astype("category")})
 .melt(id_vars="option_number")
 )
plot_df.head()
Out[20]:
<style scoped=""> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style>
option_number variable value
0 5 effectiveness_masks 1.0
1 5 effectiveness_masks 0.0
2 5 effectiveness_masks 1.0
3 5 effectiveness_masks 0.0
4 5 effectiveness_masks 0.0
In [21]:
_ = plt.figure(figsize=(15, 5))
_ = sns.stripplot(data=plot_df, 
                  x="variable", 
                  y="value", 
                  hue="option_number", 
                  dodge=True, 
                  legend=show_legend)
_ = plt.xticks(rotation=90)
No description has been provided for this image
In [58]:
DLD_df = pd.DataFrame()
for option in top_demos_list:
    tmp_df = (selection_df
    .pipe(lambda d: d[d["option_number"]==option])
    .assign(**{"combined_score": lambda d: d[feature_columns]
                .astype(str)
                .sum(axis=1)
                .astype(int)})
    .loc[:, ["option_number", "combined_score"]]
    .sort_values(by=["option_number", "combined_score"])
    .assign(**{"combined_score": lambda d: d["combined_score"].astype(str)})
    .reset_index(drop=True)
    )
    tmp_df["DLD"] = [np.nan] + [DLD(tmp_df.loc[x, "combined_score"], 
                    tmp_df.loc[x+1, "combined_score"]) for x in tmp_df.index.tolist()[:-1]]
    DLD_df = pd.concat([DLD_df, tmp_df])
In [63]:
_ = plt.figure(figsize=(5, 5))
_ = sns.histplot(data=DLD_df.dropna().assign(**{"option_number": lambda d: d["option_number"].astype(str)}), 
                 x="DLD", 
                 hue="option_number", 
                 kde=True, bins=50, 
                 legend=show_legend)
No description has been provided for this image