The choice_data object defines the choice data, it is a combination of
choice_responses and choice_covariates.
Usage
choice_data(
data_frame,
format = "wide",
column_choice = "choice",
column_decider = "deciderID",
column_occasion = NULL,
column_alternative = NULL,
column_ac_covariates = NULL,
column_as_covariates = NULL,
delimiter = "_",
cross_section = is.null(column_occasion),
choice_type = c("unordered", "ordered", "ranked")
)
generate_choice_data(
choice_effects,
choice_identifiers = generate_choice_identifiers(N = 100),
choice_covariates = generate_choice_covariates(choice_effects, choice_identifiers),
choice_parameters = generate_choice_parameters(choice_effects),
choice_preferences = generate_choice_preferences(choice_effects, choice_parameters,
choice_identifiers),
column_choice = "choice",
choice_type = c("unordered", "ordered", "ranked")
)
long_to_wide(
data_frame,
column_ac_covariates = NULL,
column_as_covariates = NULL,
column_choice = "choice",
column_alternative = "alternative",
column_decider = "deciderID",
column_occasion = NULL,
alternatives = as.character(unique(data_frame[[column_alternative]])),
delimiter = "_",
choice_type = c("unordered", "ordered", "ranked")
)
wide_to_long(
data_frame,
column_ac_covariates = NULL,
column_choice = "choice",
column_alternative = "alternative",
alternatives = NULL,
delimiter = "_",
choice_type = c("unordered", "ordered", "ranked")
)Arguments
- data_frame
[
data.frame]
Contains the choice data.- format
[
character(1)]
Format ofdata_frame. Use"wide"when each row contains all alternatives of an occasion and"long"when each row contains a single alternative.- column_choice
[
character(1)|NULL]
Column name with the observed choices. In wide layout this column should contain a single value per observation: for unordered data the value is the label of the chosen alternative, for ordered data it is the ordered factor or integer score, and for ranked data it is omitted in favor of one column per alternative (seechoice_type).In long layout the same column is evaluated once per alternative: unordered data use a binary indicator,
1orTRUEfor the chosen alternative and0orFALSEotherwise; ordered data repeats the ordinal value for every alternative, and ranked data stores consecutive ranks1:kfor the observed topkalternatives andNAfor unranked alternatives.An entirely missing response marks an occasion that is omitted from the likelihood. Set to
NULLfor purely covariate tables.- column_decider
[
character(1)]
Column name with decider identifiers.- column_occasion
[
character(1)|NULL]
Column name with occasion identifiers. Set toNULLin cross-sectional data.- column_alternative
[
character(1)|NULL]
Column name with alternative identifiers whenformat = "long".- column_ac_covariates
[
character()|NULL]
Column names with alternative-constant covariates.- column_as_covariates
[
character()|NULL]
Column names ofdata_framewith alternative-specific covariates.- delimiter
[
character(1)]
Delimiter separating alternative identifiers from covariate names whenformat = "wide".- cross_section
[
logical(1)]
Treat choice data as cross-sectional?- choice_type
[
character(1)]
Requested response type. Use"unordered"(default),"ordered", or"ranked".- choice_effects
[
choice_effects]
Achoice_effectsobject.- choice_identifiers
[
choice_identifiers]
Achoice_identifiersobject.- choice_covariates
[
choice_covariates]
Achoice_covariatesobject.- choice_parameters
[
choice_parameters]
Achoice_parametersobject.- choice_preferences
[
choice_preferences]
Achoice_preferencesobject.- alternatives
[
character(J)]
Unique labels for the choice alternatives.
Value
A choice_data tibble in the supplied format. It contains the identifier
columns, the response column(s), and the covariate columns of data_frame.
The column roles are stored as attributes, which the functions consuming
the object rely on:
formatEither
"wide"or"long".column_choiceThe name of the response column.
column_decider,column_occasionThe identifier columns.
column_alternativeThe alternative column (
format = "long").column_ac_covariates,column_as_covariatesThe names of the alternative-constant and alternative-specific covariates.
column_as_covariates_wideThe alternative-specific covariate columns in wide layout.
delimiter,cross_section,choice_typeThe corresponding input arguments.
Examples
### simulate data from a multinomial probit model
choice_effects <- choice_effects(
choice_formula = choice_formula(
formula = choice ~ A | B,
error_term = "probit",
random_effects = c("A" = "cn")
),
choice_alternatives = choice_alternatives(J = 3)
)
generate_choice_data(choice_effects = choice_effects)
#> # A tibble: 100 × 7
#> deciderID occasionID choice B A_A A_B A_C
#> * <chr> <chr> <chr> <dbl> <dbl> <dbl> <dbl>
#> 1 1 1 A 2.76 -0.0500 -0.251 0.445
#> 2 2 1 B -1.91 0.0465 0.578 0.118
#> 3 3 1 B 0.0192 0.862 -0.243 -0.206
#> 4 4 1 C 2.68 0.0296 0.550 -2.27
#> 5 5 1 B -0.665 -0.361 0.213 1.07
#> 6 6 1 C -0.976 1.11 -0.246 -1.18
#> 7 7 1 A -1.70 1.07 0.132 0.489
#> 8 8 1 A 0.237 -1.47 0.284 1.34
#> 9 9 1 B -0.110 1.32 0.524 0.607
#> 10 10 1 A 1.30 0.172 -0.0903 1.92
#> # ℹ 90 more rows
### transform from long to wide format
data("TravelMode", package = "AER")
TravelMode$choice <- TravelMode$choice == "yes"
long_to_wide(
data_frame = TravelMode,
column_alternative = "mode",
column_decider = "individual"
)
#> # A tibble: 210 × 20
#> individual income size wait_air wait_train wait_bus wait_car vcost_air
#> <fct> <int> <int> <int> <int> <int> <int> <int>
#> 1 1 35 1 69 34 35 0 59
#> 2 2 30 2 64 44 53 0 58
#> 3 3 40 1 69 34 35 0 115
#> 4 4 70 3 64 44 53 0 49
#> 5 5 45 2 64 44 53 0 60
#> 6 6 20 1 69 40 35 0 59
#> 7 7 45 1 45 34 35 0 148
#> 8 8 12 1 69 34 35 0 121
#> 9 9 40 1 69 34 35 0 59
#> 10 10 70 2 69 34 35 0 58
#> # ℹ 200 more rows
#> # ℹ 12 more variables: vcost_train <int>, vcost_bus <int>, vcost_car <int>,
#> # travel_air <int>, travel_train <int>, travel_bus <int>, travel_car <int>,
#> # gcost_air <int>, gcost_train <int>, gcost_bus <int>, gcost_car <int>,
#> # choice <fct>
### transform from wide to long format
data("Train", package = "mlogit")
wide_to_long(
data_frame = Train
)
#> # A tibble: 5,858 × 8
#> choiceid id choice alternative price time change comfort
#> <int> <int> <int> <chr> <int> <int> <int> <int>
#> 1 1 1 1 A 2400 150 0 1
#> 2 1 1 0 B 4000 150 0 1
#> 3 2 1 1 A 2400 150 0 1
#> 4 2 1 0 B 3200 130 0 1
#> 5 3 1 1 A 2400 115 0 1
#> 6 3 1 0 B 4000 115 0 0
#> 7 4 1 0 A 4000 130 0 1
#> 8 4 1 1 B 3200 150 0 0
#> 9 5 1 0 A 2400 150 0 1
#> 10 5 1 1 B 3200 150 0 0
#> # ℹ 5,848 more rows
### individual choice sets and a missing response
partial_data <- data.frame(
deciderID = c(1, 1, 2),
alternative = c("A", "B", "B"),
choice = c(1L, 0L, NA),
cost = c(1.2, 1.5, 0.8)
)
choice_data(
data_frame = partial_data,
format = "long",
column_decider = "deciderID",
column_alternative = "alternative",
column_as_covariates = "cost"
)
#> # A tibble: 3 × 4
#> deciderID alternative choice cost
#> * <dbl> <chr> <int> <dbl>
#> 1 1 A 1 1.2
#> 2 1 B 0 1.5
#> 3 2 B NA 0.8
