Skip to content

Tips for working with large datasets  #88

@ryangdar

Description

@ryangdar

Hi I'm working with a 200MB file and using the command group_similar_strings, however, this is taking so long that it's never completing (running for several days). I've tried several n_gram sizes with no luck. Do you have any tips to run on large datasets?

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions