diff options
| author | CoprDistGit <infra@openeuler.org> | 2023-04-12 07:09:16 +0000 |
|---|---|---|
| committer | CoprDistGit <infra@openeuler.org> | 2023-04-12 07:09:16 +0000 |
| commit | 88643f3aaa4db0e8cf81ee5b63887d2d293e530f (patch) | |
| tree | 0c9bdce2d734778f91da4c522a15207c0964f5d1 | |
| parent | b7343bd10b8ad4a15988e6e9a5f4a9cf3d88a56f (diff) | |
automatic import of python-minetopeneuler20.03
| -rw-r--r-- | .gitignore | 1 | ||||
| -rw-r--r-- | python-minet.spec | 514 | ||||
| -rw-r--r-- | sources | 1 |
3 files changed, 516 insertions, 0 deletions
@@ -0,0 +1 @@ +/minet-0.67.1.tar.gz diff --git a/python-minet.spec b/python-minet.spec new file mode 100644 index 0000000..6324ef4 --- /dev/null +++ b/python-minet.spec @@ -0,0 +1,514 @@ +%global _empty_manifest_terminate_build 0 +Name: python-minet +Version: 0.67.1 +Release: 1 +Summary: A webmining CLI tool & library for python. +License: MIT +URL: http://github.com/medialab/minet +Source0: https://mirrors.nju.edu.cn/pypi/web/packages/a2/80/7b6c09dad64f580c64378f90286e06df1027d849f3377f86c8aef235741f/minet-0.67.1.tar.gz +BuildArch: noarch + +Requires: python3-beautifulsoup4 +Requires: python3-browser-cookie3 +Requires: python3-casanova +Requires: python3-charset-normalizer +Requires: python3-colorama +Requires: python3-dateparser +Requires: python3-ebbe +Requires: python3-json5 +Requires: python3-keyring +Requires: python3-lxml +Requires: python3-ndjson +Requires: python3-persist-queue +Requires: python3-pyyaml +Requires: python3-quenouille +Requires: python3-soupsieve +Requires: python3-tenacity +Requires: python3-termcolor +Requires: python3-tqdm +Requires: python3-trafilatura +Requires: python3-twitwi +Requires: python3-ural +Requires: python3-urllib3 + +%description +[](https://github.com/medialab/minet/actions) [](https://zenodo.org/badge/latestdoi/169059797) [](https://pepy.tech/project/minet) + + + +**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc. + +It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work. + +In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will. + +**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md). + +## Summary + +* [What it does](#what-it-does) +* [Documented use cases](#documented-use-cases) +* [Features (from a technical standpoint)](#features-from-a-technical-standpoint) +* [Installation](#installation) +* [Upgrading](#upgrading) +* [Uninstallation](#uninstallation) +* [Documentation](#documentation) +* [Contributing](#contributing) +* [How to cite](#how-to-cite) + +## What it does + +Minet can single-handedly: +* Extract URLs from a text file (or a table) +* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff) +* Join two CSV files by matching the columns containing URLs +* From a list of URLs, resolve their redirections + * ...and check their HTTP status + * ...and download the HTML + * ...and extract hyperlinks + * ...and extract the text content and other metadata (title...) + * ...and scrape structured data (using a declarative language to define your heuristics) +* Crawl (using a declarative language to define a browsing behavior, and what to harvest) +* Mine or search: + * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess) + * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access) + * *[Mediacloud](https://mediacloud.org/)* (requires free API access) + * *[Twitter](https://twitter.com)* (requires free API access) + * *[Youtube](https://www.youtube.com/)* (requires free API access) +* Scrape (without requiring special access, often just a user account): + * *[Facebook](https://www.facebook.com/)* + * *[Instagram](https://www.instagram.com/)* + * *[Telegram](https://telegram.org/)* + * *[TikTok](https://www.tiktok.com)* + * *[Twitter](https://twitter.com)* + * *[Google Drive](https://drive.google.com)* (spreadsheets etc.) +* Grab & dump cookies from your browser +* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data + +## Documented use cases + +* [Fetching a large amount of urls](./cookbook/fetch.md) +* [Joining 2 CSV files by urls](./cookbook/url_join.md) +* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*) +* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md) +* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md) + +## Features (from a technical standpoint) + +* Multithreaded, memory-efficient fetching from the web. +* Multithreaded, scalable crawling using a comfy DSL. +* Multiprocessed raw text content extraction from HTML pages. +* Multiprocessed scraping from HTML pages using a comfy DSL. +* URL-related heuristics utilities such as extraction, normalization and matching. +* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/). + +## Installation + +**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer. + +On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it: + +```shell +sudo apt-get install curl unzip +``` + +Else, **minet** can be installed directly as a python CLI tool and library using pip: + +```shell +pip install minet +``` + +If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md). + +Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases). + +## Upgrading + +To upgrade the standalone version, simply run the install script once again: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +To upgrade the python version you can use pip thusly: + +```shell +pip install -U minet +``` + +## Uninstallation + +To uninstall the standalone version: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash +``` + +To uninstall the python version: + +```shell +pip uninstall minet +``` + +## Documentation + +* [minet as a command line tool](./docs/cli.md) +* [minet as a python library](./docs/lib.md) + +## Contributing + +To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation. + +## How to cite + +**minet** is published on [Zenodo](https://zenodo.org/) as [](https://zenodo.org/badge/latestdoi/169059797) + +You can cite it thusly: + +> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399 + + +%package -n python3-minet +Summary: A webmining CLI tool & library for python. +Provides: python-minet +BuildRequires: python3-devel +BuildRequires: python3-setuptools +BuildRequires: python3-pip +%description -n python3-minet +[](https://github.com/medialab/minet/actions) [](https://zenodo.org/badge/latestdoi/169059797) [](https://pepy.tech/project/minet) + + + +**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc. + +It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work. + +In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will. + +**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md). + +## Summary + +* [What it does](#what-it-does) +* [Documented use cases](#documented-use-cases) +* [Features (from a technical standpoint)](#features-from-a-technical-standpoint) +* [Installation](#installation) +* [Upgrading](#upgrading) +* [Uninstallation](#uninstallation) +* [Documentation](#documentation) +* [Contributing](#contributing) +* [How to cite](#how-to-cite) + +## What it does + +Minet can single-handedly: +* Extract URLs from a text file (or a table) +* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff) +* Join two CSV files by matching the columns containing URLs +* From a list of URLs, resolve their redirections + * ...and check their HTTP status + * ...and download the HTML + * ...and extract hyperlinks + * ...and extract the text content and other metadata (title...) + * ...and scrape structured data (using a declarative language to define your heuristics) +* Crawl (using a declarative language to define a browsing behavior, and what to harvest) +* Mine or search: + * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess) + * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access) + * *[Mediacloud](https://mediacloud.org/)* (requires free API access) + * *[Twitter](https://twitter.com)* (requires free API access) + * *[Youtube](https://www.youtube.com/)* (requires free API access) +* Scrape (without requiring special access, often just a user account): + * *[Facebook](https://www.facebook.com/)* + * *[Instagram](https://www.instagram.com/)* + * *[Telegram](https://telegram.org/)* + * *[TikTok](https://www.tiktok.com)* + * *[Twitter](https://twitter.com)* + * *[Google Drive](https://drive.google.com)* (spreadsheets etc.) +* Grab & dump cookies from your browser +* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data + +## Documented use cases + +* [Fetching a large amount of urls](./cookbook/fetch.md) +* [Joining 2 CSV files by urls](./cookbook/url_join.md) +* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*) +* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md) +* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md) + +## Features (from a technical standpoint) + +* Multithreaded, memory-efficient fetching from the web. +* Multithreaded, scalable crawling using a comfy DSL. +* Multiprocessed raw text content extraction from HTML pages. +* Multiprocessed scraping from HTML pages using a comfy DSL. +* URL-related heuristics utilities such as extraction, normalization and matching. +* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/). + +## Installation + +**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer. + +On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it: + +```shell +sudo apt-get install curl unzip +``` + +Else, **minet** can be installed directly as a python CLI tool and library using pip: + +```shell +pip install minet +``` + +If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md). + +Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases). + +## Upgrading + +To upgrade the standalone version, simply run the install script once again: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +To upgrade the python version you can use pip thusly: + +```shell +pip install -U minet +``` + +## Uninstallation + +To uninstall the standalone version: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash +``` + +To uninstall the python version: + +```shell +pip uninstall minet +``` + +## Documentation + +* [minet as a command line tool](./docs/cli.md) +* [minet as a python library](./docs/lib.md) + +## Contributing + +To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation. + +## How to cite + +**minet** is published on [Zenodo](https://zenodo.org/) as [](https://zenodo.org/badge/latestdoi/169059797) + +You can cite it thusly: + +> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399 + + +%package help +Summary: Development documents and examples for minet +Provides: python3-minet-doc +%description help +[](https://github.com/medialab/minet/actions) [](https://zenodo.org/badge/latestdoi/169059797) [](https://pepy.tech/project/minet) + + + +**minet** is a webmining command line tool & library for python (>= 3.7) that can be used to collect and extract data from a large variety of web sources such as raw webpages, Facebook, CrowdTangle, YouTube, Twitter, Media Cloud etc. + +It adopts a very simple approach to various webmining problems by letting you perform a variety of actions from the comfort of the command line. No database needed: raw CSV files should be sufficient to do most of the work. + +In addition, **minet** also exposes its high-level programmatic interface as a python library so you can tweak its behavior at will. + +**Shortcuts**: [Command line documentation](./docs/cli.md), [Python library documentation](./docs/lib.md). + +## Summary + +* [What it does](#what-it-does) +* [Documented use cases](#documented-use-cases) +* [Features (from a technical standpoint)](#features-from-a-technical-standpoint) +* [Installation](#installation) +* [Upgrading](#upgrading) +* [Uninstallation](#uninstallation) +* [Documentation](#documentation) +* [Contributing](#contributing) +* [How to cite](#how-to-cite) + +## What it does + +Minet can single-handedly: +* Extract URLs from a text file (or a table) +* Parse URLs (get useful information, with Facebook- and Youtube-specific stuff) +* Join two CSV files by matching the columns containing URLs +* From a list of URLs, resolve their redirections + * ...and check their HTTP status + * ...and download the HTML + * ...and extract hyperlinks + * ...and extract the text content and other metadata (title...) + * ...and scrape structured data (using a declarative language to define your heuristics) +* Crawl (using a declarative language to define a browsing behavior, and what to harvest) +* Mine or search: + * *[Buzzsumo](https://buzzsumo.com/)* (requires API acess) + * *[Crowdtangle](https://www.crowdtangle.com/)* (requires API access) + * *[Mediacloud](https://mediacloud.org/)* (requires free API access) + * *[Twitter](https://twitter.com)* (requires free API access) + * *[Youtube](https://www.youtube.com/)* (requires free API access) +* Scrape (without requiring special access, often just a user account): + * *[Facebook](https://www.facebook.com/)* + * *[Instagram](https://www.instagram.com/)* + * *[Telegram](https://telegram.org/)* + * *[TikTok](https://www.tiktok.com)* + * *[Twitter](https://twitter.com)* + * *[Google Drive](https://drive.google.com)* (spreadsheets etc.) +* Grab & dump cookies from your browser +* Dump *[Hyphe](https://hyphe.medialab.sciences-po.fr/)* data + +## Documented use cases + +* [Fetching a large amount of urls](./cookbook/fetch.md) +* [Joining 2 CSV files by urls](./cookbook/url_join.md) +* [Using minet from a Jupyter notebook](./cookbook/notebooks/Minet%20in%20a%20Jupyter%20notebook.ipynb) (*very useful to experiment with the tool or teach students*) +* [Downloading images associated with a given hashtag on Twitter](./cookbook/twitter_images.md) +* [Scraping DSL Tutorial](./cookbook/scraping_dsl.md) + +## Features (from a technical standpoint) + +* Multithreaded, memory-efficient fetching from the web. +* Multithreaded, scalable crawling using a comfy DSL. +* Multiprocessed raw text content extraction from HTML pages. +* Multiprocessed scraping from HTML pages using a comfy DSL. +* URL-related heuristics utilities such as extraction, normalization and matching. +* Data collection from various APIs such as [CrowdTangle](https://www.crowdtangle.com/). + +## Installation + +**minet** can be installed as a standalone CLI tool (currently only on mac >= 10.14, ubuntu & similar) by running the following command in your terminal: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +Don't trust us enough to pipe the result of a HTTP request into `bash`? We wouldn't either, so feel free to read the installation script [here](./scripts/install.sh) and run it on your end if you prefer. + +On ubuntu & similar you might need to install `curl` and `unzip` before running the installation script if you don't already have it: + +```shell +sudo apt-get install curl unzip +``` + +Else, **minet** can be installed directly as a python CLI tool and library using pip: + +```shell +pip install minet +``` + +If you need more help to install and use **minet** from scratch, you can check those [installation documents](./docs/install.md). + +Finally if you want to install the standalone binaries by yourself (even for windows) you can find them in each release [here](https://github.com/medialab/minet/releases). + +## Upgrading + +To upgrade the standalone version, simply run the install script once again: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/install.sh | bash +``` + +To upgrade the python version you can use pip thusly: + +```shell +pip install -U minet +``` + +## Uninstallation + +To uninstall the standalone version: + +```shell +curl -sSL https://raw.githubusercontent.com/medialab/minet/master/scripts/uninstall.sh | bash +``` + +To uninstall the python version: + +```shell +pip uninstall minet +``` + +## Documentation + +* [minet as a command line tool](./docs/cli.md) +* [minet as a python library](./docs/lib.md) + +## Contributing + +To contribute to **minet** you can check out [this](./CONTRIBUTING.md) documentation. + +## How to cite + +**minet** is published on [Zenodo](https://zenodo.org/) as [](https://zenodo.org/badge/latestdoi/169059797) + +You can cite it thusly: + +> Guillaume Plique, Pauline Breteau, Jules Farjas, Héloïse Théro, Jean Descamps, Amélie Pellé, & Laura Miguel. (2019, October 14). Minet, a webmining CLI tool & library for python. Zenodo. http://doi.org/10.5281/zenodo.4564399 + + +%prep +%autosetup -n minet-0.67.1 + +%build +%py3_build + +%install +%py3_install +install -d -m755 %{buildroot}/%{_pkgdocdir} +if [ -d doc ]; then cp -arf doc %{buildroot}/%{_pkgdocdir}; fi +if [ -d docs ]; then cp -arf docs %{buildroot}/%{_pkgdocdir}; fi +if [ -d example ]; then cp -arf example %{buildroot}/%{_pkgdocdir}; fi +if [ -d examples ]; then cp -arf examples %{buildroot}/%{_pkgdocdir}; fi +pushd %{buildroot} +if [ -d usr/lib ]; then + find usr/lib -type f -printf "/%h/%f\n" >> filelist.lst +fi +if [ -d usr/lib64 ]; then + find usr/lib64 -type f -printf "/%h/%f\n" >> filelist.lst +fi +if [ -d usr/bin ]; then + find usr/bin -type f -printf "/%h/%f\n" >> filelist.lst +fi +if [ -d usr/sbin ]; then + find usr/sbin -type f -printf "/%h/%f\n" >> filelist.lst +fi +touch doclist.lst +if [ -d usr/share/man ]; then + find usr/share/man -type f -printf "/%h/%f.gz\n" >> doclist.lst +fi +popd +mv %{buildroot}/filelist.lst . +mv %{buildroot}/doclist.lst . + +%files -n python3-minet -f filelist.lst +%dir %{python3_sitelib}/* + +%files help -f doclist.lst +%{_docdir}/* + +%changelog +* Wed Apr 12 2023 Python_Bot <Python_Bot@openeuler.org> - 0.67.1-1 +- Package Spec generated @@ -0,0 +1 @@ +01ac08e793d222cda6551ce67a595003 minet-0.67.1.tar.gz |
