scikit-learn-contrib/hdbscan

Introduce sample weighting in HDBSCAN

Open

#148 opened on Dec 5, 2017

 (18 comments) (10 reactions) (0 assignees)Jupyter Notebook (535 forks)github user discovery
help wantednew feature

Repository metrics

Stars
 (3,137 stars)
PR merge metrics
 (No merged PRs in 30d)

Description

Like for DBSCAN it would be very useful to also have the possibility to use a "sample_weight" parameter when performing HDBSCAN clustering trough the fit method. This to take into account possible presence of data containing element duplicates.

DBSCAN is just using sample_weight in the following way:

n_neighbors = np.array([np.sum(sample_weight[neighbors]) for neighbors in neighborhoods])

do you know if this is already planned at same time or not?

Contributor guide