summaryrefslogtreecommitdiff
diff options
context:
space:
mode:
-rw-r--r--.gitignore1
-rw-r--r--python-minet.spec514
-rw-r--r--sources1
3 files changed, 516 insertions, 0 deletions
diff --git a/.gitignore b/.gitignore
index e69de29..1ff7bb2 100644
--- a/.gitignore
+++ b/.gitignore
@@ -0,0 +1 @@
+/minet-0.67.1.tar.gz
diff --git a/python-minet.spec b/python-minet.spec
new file mode 100644
index 0000000..6324ef4
--- /dev/null
+++ b/python-minet.spec
@@ -0,0 +1,514 @@
+%global _empty_manifest_terminate_build 0
+Name: python-minet
+Version: 0.67.1
+Release: 1
+Summary: A webmining CLI tool & library for python.
+License: MIT
+URL: http://github.com/medialab/minet
+Source0: https://mirrors.nju.edu.cn/pypi/web/packages/a2/80/7b6c09dad64f580c64378f90286e06df1027d849f3377f86c8aef235741f/minet-0.67.1.tar.gz
+BuildArch: noarch
+
+Requires: python3-beautifulsoup4
+Requires: python3-browser-cookie3
+Requires: python3-casanova
+Requires: python3-charset-normalizer
+Requires: python3-colorama
+Requires: python3-dateparser
+Requires: python3-ebbe
+Requires: python3-json5
+Requires: python3-keyring
+Requires: python3-lxml
+Requires: python3-ndjson
+Requires: python3-persist-queue
+Requires: python3-pyyaml
+Requires: python3-quenouille
+Requires: python3-soupsieve
+Requires: python3-tenacity
+Requires: python3-termcolor
+Requires: python3-tqdm
+Requires: python3-trafilatura
+Requires: python3-twitwi
+Requires: python3-ural
+Requires: python3-urllib3
+
+%description
+[![Build Status](https://github.com/medialab/minet/workflows/Tests/badge.svg)](https://github.com/medialab/minet/actions) [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797) [![download number](https://static.pepy.tech/badge/minet)](https://pepy.tech/project/minet)
+
+![Minet](img/minet.png)
+
+**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc.
+
+It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work.
+
+In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will.
+
+**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md).
+
+## Summary
+
+* [What it does](#what-it-does)
+* [Documented use cases](#documented-use-cases)
+* [Features (from a technical standpoint)](#features-from-a-technical-standpoint)
+* [Installation](#installation)
+* [Upgrading](#upgrading)
+* [Uninstallation](#uninstallation)
+* [Documentation](#documentation)
+* [Contributing](#contributing)
+* [How to cite](#how-to-cite)
+
+## What it does
+
+Minet can single-handedly:
+* Extract URLs from a text file (or a table)
+* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff)
+* Join two CSV files by matching the columns containing URLs
+* From a list of URLs, resolve their redirections
+ * ...and check their HTTP status
+ * ...and download the HTML
+ * ...and extract hyperlinks
+ * ...and extract the text content and other metadata (title...)
+ * ...and scrape structured data (using a declarative language to define your heuristics)
+* Crawl (using a declarative language to define a browsing behavior, and what to harvest)
+* Mine or search:
+ * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess)
+ * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access)
+ * *[Mediacloud](https://mediacloud.org/)* (requires free API access)
+ * *[Twitter](https://twitter.com)* (requires free API access)
+ * *[Youtube](https://www.youtube.com/)* (requires free API access)
+* Scrape (without requiring special access, often just a user account):
+ * *[Facebook](https://www.facebook.com/)*
+ * *[Instagram](https://www.instagram.com/)*
+ * *[Telegram](https://telegram.org/)*
+ * *[TikTok](https://www.tiktok.com)*
+ * *[Twitter](https://twitter.com)*
+ * *[Google Drive](https://drive.google.com)* (spreadsheets etc.)
+* Grab & dump cookies from your browser
+* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data
+
+## Documented use cases
+
+* [Fetching a large amount of urls](./cookbook/fetch.md)
+* [Joining 2 CSV files by urls](./cookbook/url_join.md)
+* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*)
+* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md)
+* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md)
+
+## Features (from a technical standpoint)
+
+* Multithreaded, memory-efficient fetching from the web.
+* Multithreaded, scalable crawling using a comfy DSL.
+* Multiprocessed raw text content extraction from HTML pages.
+* Multiprocessed scraping from HTML pages using a comfy DSL.
+* URL-related heuristics utilities such as extraction, normalization and matching.
+* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/).
+
+## Installation
+
+**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer.
+
+On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it:
+
+```shell
+sudo apt-get install curl unzip
+```
+
+Else, **minet** can be installed directly as a python CLI tool and library using pip:
+
+```shell
+pip install minet
+```
+
+If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md).
+
+Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases).
+
+## Upgrading
+
+To upgrade the standalone version, simply run the install script once again:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+To upgrade the python version you can use pip thusly:
+
+```shell
+pip install -U minet
+```
+
+## Uninstallation
+
+To uninstall the standalone version:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash
+```
+
+To uninstall the python version:
+
+```shell
+pip uninstall minet
+```
+
+## Documentation
+
+* [minet as a command line tool](./docs/cli.md)
+* [minet as a python library](./docs/lib.md)
+
+## Contributing
+
+To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation.
+
+## How to cite
+
+**minet** is published on [Zenodo](https://zenodo.org/) as [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797)
+
+You can cite it thusly:
+
+> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399
+
+
+%package -n python3-minet
+Summary: A webmining CLI tool & library for python.
+Provides: python-minet
+BuildRequires: python3-devel
+BuildRequires: python3-setuptools
+BuildRequires: python3-pip
+%description -n python3-minet
+[![Build Status](https://github.com/medialab/minet/workflows/Tests/badge.svg)](https://github.com/medialab/minet/actions) [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797) [![download number](https://static.pepy.tech/badge/minet)](https://pepy.tech/project/minet)
+
+![Minet](img/minet.png)
+
+**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc.
+
+It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work.
+
+In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will.
+
+**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md).
+
+## Summary
+
+* [What it does](#what-it-does)
+* [Documented use cases](#documented-use-cases)
+* [Features (from a technical standpoint)](#features-from-a-technical-standpoint)
+* [Installation](#installation)
+* [Upgrading](#upgrading)
+* [Uninstallation](#uninstallation)
+* [Documentation](#documentation)
+* [Contributing](#contributing)
+* [How to cite](#how-to-cite)
+
+## What it does
+
+Minet can single-handedly:
+* Extract URLs from a text file (or a table)
+* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff)
+* Join two CSV files by matching the columns containing URLs
+* From a list of URLs, resolve their redirections
+ * ...and check their HTTP status
+ * ...and download the HTML
+ * ...and extract hyperlinks
+ * ...and extract the text content and other metadata (title...)
+ * ...and scrape structured data (using a declarative language to define your heuristics)
+* Crawl (using a declarative language to define a browsing behavior, and what to harvest)
+* Mine or search:
+ * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess)
+ * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access)
+ * *[Mediacloud](https://mediacloud.org/)* (requires free API access)
+ * *[Twitter](https://twitter.com)* (requires free API access)
+ * *[Youtube](https://www.youtube.com/)* (requires free API access)
+* Scrape (without requiring special access, often just a user account):
+ * *[Facebook](https://www.facebook.com/)*
+ * *[Instagram](https://www.instagram.com/)*
+ * *[Telegram](https://telegram.org/)*
+ * *[TikTok](https://www.tiktok.com)*
+ * *[Twitter](https://twitter.com)*
+ * *[Google Drive](https://drive.google.com)* (spreadsheets etc.)
+* Grab & dump cookies from your browser
+* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data
+
+## Documented use cases
+
+* [Fetching a large amount of urls](./cookbook/fetch.md)
+* [Joining 2 CSV files by urls](./cookbook/url_join.md)
+* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*)
+* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md)
+* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md)
+
+## Features (from a technical standpoint)
+
+* Multithreaded, memory-efficient fetching from the web.
+* Multithreaded, scalable crawling using a comfy DSL.
+* Multiprocessed raw text content extraction from HTML pages.
+* Multiprocessed scraping from HTML pages using a comfy DSL.
+* URL-related heuristics utilities such as extraction, normalization and matching.
+* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/).
+
+## Installation
+
+**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer.
+
+On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it:
+
+```shell
+sudo apt-get install curl unzip
+```
+
+Else, **minet** can be installed directly as a python CLI tool and library using pip:
+
+```shell
+pip install minet
+```
+
+If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md).
+
+Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases).
+
+## Upgrading
+
+To upgrade the standalone version, simply run the install script once again:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+To upgrade the python version you can use pip thusly:
+
+```shell
+pip install -U minet
+```
+
+## Uninstallation
+
+To uninstall the standalone version:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash
+```
+
+To uninstall the python version:
+
+```shell
+pip uninstall minet
+```
+
+## Documentation
+
+* [minet as a command line tool](./docs/cli.md)
+* [minet as a python library](./docs/lib.md)
+
+## Contributing
+
+To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation.
+
+## How to cite
+
+**minet** is published on [Zenodo](https://zenodo.org/) as [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797)
+
+You can cite it thusly:
+
+> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399
+
+
+%package help
+Summary: Development documents and examples for minet
+Provides: python3-minet-doc
+%description help
+[![Build Status](https://github.com/medialab/minet/workflows/Tests/badge.svg)](https://github.com/medialab/minet/actions) [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797) [![download number](https://static.pepy.tech/badge/minet)](https://pepy.tech/project/minet)
+
+![Minet](img/minet.png)
+
+**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc.
+
+It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work.
+
+In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will.
+
+**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md).
+
+## Summary
+
+* [What it does](#what-it-does)
+* [Documented use cases](#documented-use-cases)
+* [Features (from a technical standpoint)](#features-from-a-technical-standpoint)
+* [Installation](#installation)
+* [Upgrading](#upgrading)
+* [Uninstallation](#uninstallation)
+* [Documentation](#documentation)
+* [Contributing](#contributing)
+* [How to cite](#how-to-cite)
+
+## What it does
+
+Minet can single-handedly:
+* Extract URLs from a text file (or a table)
+* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff)
+* Join two CSV files by matching the columns containing URLs
+* From a list of URLs, resolve their redirections
+ * ...and check their HTTP status
+ * ...and download the HTML
+ * ...and extract hyperlinks
+ * ...and extract the text content and other metadata (title...)
+ * ...and scrape structured data (using a declarative language to define your heuristics)
+* Crawl (using a declarative language to define a browsing behavior, and what to harvest)
+* Mine or search:
+ * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess)
+ * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access)
+ * *[Mediacloud](https://mediacloud.org/)* (requires free API access)
+ * *[Twitter](https://twitter.com)* (requires free API access)
+ * *[Youtube](https://www.youtube.com/)* (requires free API access)
+* Scrape (without requiring special access, often just a user account):
+ * *[Facebook](https://www.facebook.com/)*
+ * *[Instagram](https://www.instagram.com/)*
+ * *[Telegram](https://telegram.org/)*
+ * *[TikTok](https://www.tiktok.com)*
+ * *[Twitter](https://twitter.com)*
+ * *[Google Drive](https://drive.google.com)* (spreadsheets etc.)
+* Grab & dump cookies from your browser
+* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data
+
+## Documented use cases
+
+* [Fetching a large amount of urls](./cookbook/fetch.md)
+* [Joining 2 CSV files by urls](./cookbook/url_join.md)
+* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*)
+* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md)
+* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md)
+
+## Features (from a technical standpoint)
+
+* Multithreaded, memory-efficient fetching from the web.
+* Multithreaded, scalable crawling using a comfy DSL.
+* Multiprocessed raw text content extraction from HTML pages.
+* Multiprocessed scraping from HTML pages using a comfy DSL.
+* URL-related heuristics utilities such as extraction, normalization and matching.
+* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/).
+
+## Installation
+
+**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer.
+
+On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it:
+
+```shell
+sudo apt-get install curl unzip
+```
+
+Else, **minet** can be installed directly as a python CLI tool and library using pip:
+
+```shell
+pip install minet
+```
+
+If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md).
+
+Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases).
+
+## Upgrading
+
+To upgrade the standalone version, simply run the install script once again:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash
+```
+
+To upgrade the python version you can use pip thusly:
+
+```shell
+pip install -U minet
+```
+
+## Uninstallation
+
+To uninstall the standalone version:
+
+```shell
+curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash
+```
+
+To uninstall the python version:
+
+```shell
+pip uninstall minet
+```
+
+## Documentation
+
+* [minet as a command line tool](./docs/cli.md)
+* [minet as a python library](./docs/lib.md)
+
+## Contributing
+
+To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation.
+
+## How to cite
+
+**minet** is published on [Zenodo](https://zenodo.org/) as [![DOI](https://zenodo.org/badge/169059797.svg)](https://zenodo.org/badge/latestdoi/169059797)
+
+You can cite it thusly:
+
+> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399
+
+
+%prep
+%autosetup -n minet-0.67.1
+
+%build
+%py3_build
+
+%install
+%py3_install
+install -d -m755 %{buildroot}/%{_pkgdocdir}
+if [ -d doc ]; then cp -arf doc %{buildroot}/%{_pkgdocdir}; fi
+if [ -d docs ]; then cp -arf docs %{buildroot}/%{_pkgdocdir}; fi
+if [ -d example ]; then cp -arf example %{buildroot}/%{_pkgdocdir}; fi
+if [ -d examples ]; then cp -arf examples %{buildroot}/%{_pkgdocdir}; fi
+pushd %{buildroot}
+if [ -d usr/lib ]; then
+ find usr/lib -type f -printf "/%h/%f\n" >> filelist.lst
+fi
+if [ -d usr/lib64 ]; then
+ find usr/lib64 -type f -printf "/%h/%f\n" >> filelist.lst
+fi
+if [ -d usr/bin ]; then
+ find usr/bin -type f -printf "/%h/%f\n" >> filelist.lst
+fi
+if [ -d usr/sbin ]; then
+ find usr/sbin -type f -printf "/%h/%f\n" >> filelist.lst
+fi
+touch doclist.lst
+if [ -d usr/share/man ]; then
+ find usr/share/man -type f -printf "/%h/%f.gz\n" >> doclist.lst
+fi
+popd
+mv %{buildroot}/filelist.lst .
+mv %{buildroot}/doclist.lst .
+
+%files -n python3-minet -f filelist.lst
+%dir %{python3_sitelib}/*
+
+%files help -f doclist.lst
+%{_docdir}/*
+
+%changelog
+* Wed Apr 12 2023 Python_Bot <Python_Bot@openeuler.org> - 0.67.1-1
+- Package Spec generated
diff --git a/sources b/sources
new file mode 100644
index 0000000..b8706f1
--- /dev/null
+++ b/sources
@@ -0,0 +1 @@
+01ac08e793d222cda6551ce67a595003 minet-0.67.1.tar.gz