Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Inquiry about Web Pipeline Availability #151

Open
codefly13 opened this issue Apr 22, 2024 · 2 comments
Open

Inquiry about Web Pipeline Availability #151

codefly13 opened this issue Apr 22, 2024 · 2 comments

Comments

@codefly13
Copy link

I hope you are doing well. I came across a reference to the "Web Pipeline" in the paper "Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research" and I am very interested in exploring it further. However, it seems that the pipeline is still in preparation. I would like to kindly inquire about the availability of the "Web Pipeline". Is there any information on when it might be released for public use?

@codefly13 codefly13 changed the title Inquiry about CommonCrawl WARC Pipeline Availability Inquiry about Web Pipeline Availability Apr 22, 2024
@dumitrac
Copy link

Hi @codefly13 - all of it is already available in the dolma toolkit (i.e. this repo).
Please let me know if you're looking for something different.

@OxxoCodes
Copy link

@dumitrac I'm interested in this as well. I'd like to utilize the Dolma toolkit to perform some filtering on CC data (which is what I assume @codefly13 was attempting to perform as well). However, I don't see an example of how to do this in the repo, and the following pipeline is just marked as being WIP: https://github.com/allenai/dolma/tree/main/sources/cc_warc

I'm very new to Dolma so there's a good chance I'm just missing something. Would appreciate some pointers. Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
None yet
Projects
None yet
Development

No branches or pull requests

3 participants