Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Token filtering with a measure of association? #270

Open
mfra123 opened this issue Aug 29, 2024 · 0 comments
Open

Token filtering with a measure of association? #270

mfra123 opened this issue Aug 29, 2024 · 0 comments

Comments

@mfra123
Copy link

mfra123 commented Aug 29, 2024

Would it be possible (and useful) to add a functionality to filter token lists by a measure of association with the outcome e.g. a chi squared / log odds?

When working with heavily imbalanced data selecting top n tokens (or similar) with step_tokenfilter is unlikely to include discriminative features, unless the number of terms included is very large or the training data is artificially balanced before token filtering. Similarly if you wish to include a mixture of unigrams and bigrams etc less frequent terms which are discriminative are unlikely to be included unless large numbers of tokens are included, many of which will be common words often with little predictive value.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

No branches or pull requests

1 participant