irspack.recommenders.AsymmetricCosineUserKNNRecommender#
- class irspack.recommenders.AsymmetricCosineUserKNNRecommender(X_train_all, shrinkage=0.0, alpha=0.5, top_k=100, feature_weighting='NONE', bm25_k1=1.2, bm25_b=0.75, n_threads=None)[source]#
Bases:
BaseUserKNNRecommenderK-nearest neighbor recommender system based on asymmetric cosine similarity. That is, the similarity matrix
Uis given by (row-wise top-k restricted)\[\mathrm{U}_{u,v} = \frac{\sum_{i} X_{ui} X_{vi}}{||X_{u*}||^{2\alpha}_2 ||X_{v*}||^{2(1-\alpha)}_2 + \mathrm{shrinkage}}\]- Parameters:
X_train_all (Union[scipy.sparse.csr_matrix, scipy.sparse.csc_matrix]) – Input interaction matrix.
shrinkage (float, optional) – The shrinkage parameter for regularization. Defaults to 0.0.
alpha (bool, optional) – Specifies \(\alpha\). Defaults to 0.5.
top_k (int, optional) – Specifies the maximal number of allowed neighbors. Defaults to 100.
feature_weighting (str, optional) –
Specifies how to weight the feature. Must be one of:
”NONE” : no feature weighting
”TF_IDF” : TF-IDF weighting
”BM_25” : Okapi BM-25 weighting
Defaults to “NONE”.
bm25_k1 (float, optional) – The k1 parameter for BM25. Ignored if
feature_weightingis not “BM_25”. Defaults to 1.2.bm25_b (float, optional) – The b parameter for BM25. Ignored if
feature_weightingis not “BM_25”. Defaults to 0.75.n_threads (Optional[int], optional) – Specifies the number of threads to use for the computation. If
None, the environment variable"IRSPACK_NUM_THREADS_DEFAULT"will be looked up, and if the variable is not set, it will be set toos.cpu_count(). Defaults to None.
- __init__(X_train_all, shrinkage=0.0, alpha=0.5, top_k=100, feature_weighting='NONE', bm25_k1=1.2, bm25_b=0.75, n_threads=None)[source]#
- Parameters:
X_train_all (csr_matrix | csc_matrix)
shrinkage (float)
alpha (float)
top_k (int)
feature_weighting (str)
bm25_k1 (float)
bm25_b (float)
n_threads (int | None)
Methods
__init__(X_train_all[, shrinkage, alpha, ...])default_suggest_parameter(trial, ...)from_config(X_train_all, config)get_score(user_indices)Compute the item recommendation score for a subset of users.
get_score_block(begin, end)Compute the score for a block of the users.
Compute the item recommendation score for unseen users whose profiles are given as another user-item relation matrix.
Compute the item recommendation score for unseen users whose profiles are given as another user-item relation matrix.
get_score_remove_seen(user_indices)Compute the item score and mask the item in the training set.
get_score_remove_seen_block(begin, end)Compute the score for a block of the users, and mask the items in the training set.
learn()Learns and returns itself.
learn_with_optimizer(evaluator, trial[, ...])Learning procedures with early stopping and pruning.
tune(data, evaluator[, study, n_trials, ...])Perform the optimization step.
Attributes
The computed user-user similarity weight matrix.
default_tune_rangeU_The matrix to feed into recommender.
- property U: csr_matrix | csc_matrix | ndarray#
The computed user-user similarity weight matrix.
- X_train_all: sps.csr_matrix#
The matrix to feed into recommender.
- config_class#
alias of
AsymmetricCosineUserKNNConfig
- get_score(user_indices)#
Compute the item recommendation score for a subset of users.
- Parameters:
user_indices (ndarray) – The index defines the subset of users.
- Returns:
The item scores. Its shape will be (len(user_indices), self.n_items)
- Return type:
ndarray
- get_score_block(begin, end)#
Compute the score for a block of the users.
- Parameters:
begin (int) – where the evaluated user block begins.
end (int) – where the evaluated user block ends.
- Returns:
The item scores. Its shape will be (end - begin, self.n_items)
- Return type:
ndarray
- get_score_cold_user(X)#
Compute the item recommendation score for unseen users whose profiles are given as another user-item relation matrix.
- Parameters:
X (csr_matrix | csc_matrix) – The profile user-item relation matrix for unseen users. Its number of rows is arbitrary, but the number of columns must be self.n_items.
- Returns:
Computed item scores for users. Its shape is equal to X.
- Return type:
ndarray
- get_score_cold_user_remove_seen(X)#
Compute the item recommendation score for unseen users whose profiles are given as another user-item relation matrix. The score will then be masked by the input.
- Parameters:
X (csr_matrix | csc_matrix) – The profile user-item relation matrix for unseen users. Its number of rows is arbitrary, but the number of columns must be self.n_items.
- Returns:
Computed & masked item scores for users. Its shape is equal to X.
- Return type:
ndarray
- get_score_remove_seen(user_indices)#
Compute the item score and mask the item in the training set. Masked items will have the score -inf.
- Parameters:
user_indices (ndarray) – Specifies the subset of users.
- Returns:
The masked item scores. Its shape will be (len(user_indices), self.n_items)
- Return type:
ndarray
- get_score_remove_seen_block(begin, end)#
Compute the score for a block of the users, and mask the items in the training set. Masked items will have the score -inf.
- Parameters:
begin (int) – where the evaluated user block begins.
end (int) – where the evaluated user block ends.
- Returns:
The masked item scores. Its shape will be (end - begin, self.n_items)
- Return type:
ndarray
- learn()#
Learns and returns itself.
- Returns:
The model after fitting process.
- Parameters:
self (R)
- Return type:
R
- learn_with_optimizer(evaluator, trial, max_epoch=128, validate_epoch=5, score_degradation_max=5)#
Learning procedures with early stopping and pruning.
- Parameters:
evaluator (ForwardRef('evaluation.Evaluator') | None) – The evaluator to measure the score.
trial (ForwardRef('Trial') | None) – The current optuna trial under the study (if any.)
max_epoch (int) – Maximal number of epochs. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 128.
validate_epoch (int) – The frequency of validation score measurement. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 5.
validate_epoch – The frequency of validation score measurement. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 5.
score_degradation_max (int) – Maximal number of allowed score degradation. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 5.
- Return type:
None
- classmethod tune(data, evaluator, study=None, n_trials=20, timeout=None, data_suggest_function=None, parameter_suggest_function=None, tuning_random_seed=None, prunning_n_startup_trials=10, max_epoch=128, validate_epoch=5, score_degradation_max=5, logger=None, **recommender_params)#
Perform the optimization step. An optuna.Study object can be supplied or created inside this function.
- Parameters:
data (csr_matrix | csc_matrix | None) – The training data. You can also provide tunable parameter dependent training data by providing data_suggest_function. In that case, data must be None.
evaluator (evaluation.Evaluator) – The validation evaluator that measures the performance of the recommenders.
study (ForwardRef('Study') | None) – An existing Optuna study. If
None, a new study is created.n_trials (int) – The number of expected trials (including pruned ones). Defaults to 20.
timeout (int | None) – If set to some value (in seconds), the study will exit after that time period. Note that the running trials is not interrupted, though. Defaults to None.
data_suggest_function (Callable[[ForwardRef('Trial')], csr_matrix | csc_matrix] | None) – If not None, this must be a function which takes optuna.Trial as its argument and returns training data. Defaults to None.
parameter_suggest_function (Callable[[ForwardRef('Trial')], Dict[str, Any]] | None) – If not None, this must be a function which takes optuna.Trial as its argument and returns Dict[str, Any] (i.e., some keyword arguments of the recommender class). If None, cls.default_suggest_parameter will be used for the parameter suggestion. Defaults to None.
tuning_random_seed (int | None) – The random seed to control optuna.samplers.TPESampler. Defaults to None. Ignored when
studyis provided.prunning_n_startup_trials (int) – n_startup_trials argument passed to the constructor of optuna.pruners.MedianPruner.
max_epoch (int) – The maximal number of epochs for the training. If iterative learning procedure is not available, this parameter will be ignored.
validate_epoch (int, optional) – The frequency of validation score measurement. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 5.
score_degradation_max (int, optional) – Maximal number of allowed score degradation. If iterative learning procedure is not available, this parameter will be ignored. Defaults to 5. Defaults to 5.
**recommender_params (Any) – Fixed keyword arguments passed to the recommender constructor for every trial. These override suggested parameters with the same name, and are not included in the returned best parameters.
logger (Logger | None)
**recommender_params
- Returns:
A tuple that consists of
A dict containing the best suggested and learnt parameters. Fixed
recommender_paramsare not included.A
pandas.DataFramethat contains the history of optimization.
- Return type:
Tuple[Dict[str, Any], DataFrame]