330 lines
8.9 KiB
Markdown
330 lines
8.9 KiB
Markdown
# rule34vault downloader ideas
|
|
|
|
## User requirements / comments to keep
|
|
|
|
- Download bookmarks for a public username; no authorization required.
|
|
- Bookmark pages are available as `https://rule34vault.com/u/<username>?page=<n>`.
|
|
- Only username should be required for normal use.
|
|
- Use browser-like headers to reduce chance of request blocking/blocklisting.
|
|
- Do **not** rely on `HEAD` only. Some servers may reject or mis-handle `HEAD`; use `GET` with streaming/range checks when needed.
|
|
- For normal image posts, first try predictable original-file URLs:
|
|
- `https://rule34vault.com/posts/<prefix>/<post_id>/<post_id>.jpg`
|
|
- `https://r34xyz.b-cdn.net/posts/<prefix>/<post_id>/<post_id>.jpg`
|
|
- `<prefix>` is the first digits group from the post id path. Example: post `1263322` uses prefix `1263`.
|
|
- Some videos require opening the post page and extracting real `<video>/<source src=...>` URLs from the rendered HTML.
|
|
- Create a plain text list of downloaded files in the same order they appear in bookmarks.
|
|
- Plain txt state/log files are enough. Do **not** use SQLite.
|
|
- Use multiple tries and exponential sleep after failures.
|
|
- Project must use a new venv.
|
|
|
|
## Real site structure inspected
|
|
|
|
Inspected with browser-like headers from the project venv.
|
|
|
|
### Video post example
|
|
|
|
URL:
|
|
|
|
```text
|
|
https://rule34vault.com/post/738794
|
|
```
|
|
|
|
The page contains a real video element in returned HTML:
|
|
|
|
```html
|
|
<video controls loop class="video" style="aspect-ratio: 1280/720;">
|
|
<source type="video/mp4" src="https://r34xyz.b-cdn.net/posts/738/738794/738794.480.mp4">
|
|
</video>
|
|
```
|
|
|
|
So video extraction should parse the post HTML with BeautifulSoup and collect:
|
|
|
|
1. `video[src]`
|
|
2. `video source[src]`
|
|
3. fallback: any `source[src]` where `type` starts with `video/`
|
|
|
|
For post `996379`, real page had multiple video sources:
|
|
|
|
```text
|
|
https://r34xyz.b-cdn.net/posts/996/996379/996379.720.hevc.mp4
|
|
https://r34xyz.b-cdn.net/posts/996/996379/996379.480.mp4
|
|
```
|
|
|
|
Prefer broadly-compatible sources first, e.g. non-HEVC MP4 before HEVC if both exist.
|
|
|
|
### Image post examples
|
|
|
|
For:
|
|
|
|
```text
|
|
https://rule34vault.com/post/1263322
|
|
```
|
|
|
|
The page HTML contains a small display image:
|
|
|
|
```html
|
|
<img class="img" src="/posts/1263/1263322/1263322.small.jpg">
|
|
```
|
|
|
|
But the original image is available directly:
|
|
|
|
```text
|
|
https://rule34vault.com/posts/1263/1263322/1263322.jpg
|
|
https://r34xyz.b-cdn.net/posts/1263/1263322/1263322.jpg
|
|
```
|
|
|
|
Both returned `200 image/jpeg` during inspection.
|
|
|
|
Another inspected example:
|
|
|
|
```text
|
|
https://rule34vault.com/post/1263303
|
|
```
|
|
|
|
Sometimes the post page `img[src]` points only to a small/preview file:
|
|
|
|
```text
|
|
https://rule34vault.com/posts/1263/1263303/1263303.small.jpg
|
|
```
|
|
|
|
In this case the downloader must normalize the source URL to the original filename by removing the size suffix before the extension:
|
|
|
|
```text
|
|
https://rule34vault.com/posts/1263/1263303/1263303.jpg
|
|
```
|
|
|
|
This URL returned `200 image/jpeg` during inspection. Do not download `.small.jpg` when the corresponding original image exists.
|
|
|
|
### Bookmark page structure
|
|
|
|
On:
|
|
|
|
```text
|
|
https://rule34vault.com/u/krosh?page=6
|
|
```
|
|
|
|
Post links were **not** regular `<a href="/post/...">` links in the DOM. The bookmark data was inside:
|
|
|
|
```html
|
|
<script id="ng-state" type="application/json">...</script>
|
|
```
|
|
|
|
Relevant key looked like:
|
|
|
|
```json
|
|
"post:/api/v2/post/search/bookmarked/50311": {
|
|
"items": [
|
|
{"id": 1075647, "type": 0, ...},
|
|
{"id": 1081537, "type": 0, ...},
|
|
...
|
|
]
|
|
}
|
|
```
|
|
|
|
Therefore bookmark extraction should prefer parsing `script#ng-state` JSON and reading `items` from the key starting with:
|
|
|
|
```text
|
|
post:/api/v2/post/search/bookmarked/
|
|
```
|
|
|
|
Keep item order exactly as found in that `items` array.
|
|
|
|
The API also works directly as POST:
|
|
|
|
```text
|
|
POST https://rule34vault.com/api/v2/post/search/bookmarked/<user_id>
|
|
```
|
|
|
|
Examples observed:
|
|
|
|
- Body `{"skip": 100, "take": 20}` returned 20 bookmark items.
|
|
- Body `{"take": 20, "page": 6}` returned page-like bookmark items.
|
|
|
|
Simplest robust path:
|
|
|
|
1. Fetch profile page `/u/<username>?page=<page>`.
|
|
2. Parse `ng-state`.
|
|
3. Get user id from key `get:/api/v2/account/user/<username>`.
|
|
4. Get bookmark items from `post:/api/v2/post/search/bookmarked/<user_id>`.
|
|
5. Optionally use the direct POST API for pagination once user id is known.
|
|
|
|
## Headers
|
|
|
|
Use one shared `requests.Session()` with realistic headers:
|
|
|
|
```python
|
|
HEADERS = {
|
|
"User-Agent": "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 "
|
|
"(KHTML, like Gecko) Chrome/126.0 Safari/537.36",
|
|
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
|
|
"Accept-Language": "en-US,en;q=0.9",
|
|
"Connection": "keep-alive",
|
|
"Referer": "https://rule34vault.com/",
|
|
}
|
|
```
|
|
|
|
For direct API POST, use:
|
|
|
|
```python
|
|
API_HEADERS = {
|
|
**HEADERS,
|
|
"Accept": "application/json,text/plain,*/*",
|
|
"Origin": "https://rule34vault.com",
|
|
"Content-Type": "application/json",
|
|
}
|
|
```
|
|
|
|
For media download, set referer to the post page:
|
|
|
|
```python
|
|
headers = {**HEADERS, "Referer": f"https://rule34vault.com/post/{post_id}"}
|
|
```
|
|
|
|
## Retry policy
|
|
|
|
Every network operation should use retry wrapper:
|
|
|
|
```python
|
|
for attempt in range(max_tries):
|
|
try:
|
|
...
|
|
except requests.RequestException:
|
|
sleep = base_delay * (2 ** attempt)
|
|
time.sleep(sleep)
|
|
```
|
|
|
|
Also retry on temporary HTTP statuses:
|
|
|
|
```text
|
|
429, 500, 502, 503, 504
|
|
```
|
|
|
|
Do not retry permanent `404` for candidate media URLs; just try the next candidate.
|
|
|
|
## Download order log
|
|
|
|
Create a txt manifest in download order:
|
|
|
|
```text
|
|
download_order.txt
|
|
```
|
|
|
|
Append one line per successfully downloaded post, exactly in bookmark order:
|
|
|
|
```text
|
|
000001 1075647 downloads/1075/1075647.jpg https://rule34vault.com/posts/1075/1075647/1075647.jpg
|
|
000002 1081537 downloads/1081/1081537.jpg https://r34xyz.b-cdn.net/posts/1081/1081537/1081537.jpg
|
|
000003 996379 downloads/996/996379.mp4 https://r34xyz.b-cdn.net/posts/996/996379/996379.480.mp4
|
|
```
|
|
|
|
Also keep simple txt logs:
|
|
|
|
```text
|
|
seen.txt # post ids already processed successfully
|
|
failed.txt # post id + reason
|
|
skipped.txt # already existing files
|
|
```
|
|
|
|
No SQLite.
|
|
|
|
## Media resolution strategy
|
|
|
|
For each bookmark item in order:
|
|
|
|
1. If already downloaded, skip and log to `skipped.txt`.
|
|
2. If `type == 0` image:
|
|
- Try original candidate URLs by extension and host.
|
|
- Use streaming `GET`, not only `HEAD`.
|
|
- Validate `status_code == 200` and `Content-Type` starts with `image/`.
|
|
3. If candidates fail, fetch post page and parse actual HTML.
|
|
- For image posts, inspect `img[src]` values from the real post page.
|
|
- If the only discovered image is a display derivative such as `<post_id>.small.jpg`, `<post_id>.preview.jpg`, `<post_id>.thumbnail.jpg`, etc., derive and try the original URL without the size suffix before downloading:
|
|
- `1263303.small.jpg` -> `1263303.jpg`
|
|
- `1263303.preview.jpg` -> `1263303.jpg`
|
|
- Download the higher-resolution original when it returns a valid image response; use the derivative only as a last fallback if the original does not exist.
|
|
4. For videos (`type == 1`) or if image candidates fail:
|
|
- Fetch `https://rule34vault.com/post/<post_id>`.
|
|
- Parse `<video>` and `<source>` tags.
|
|
- Choose best compatible source.
|
|
5. Download media to grouped directory:
|
|
|
|
```text
|
|
downloads/<prefix>/<post_id>.<ext>
|
|
```
|
|
|
|
## Candidate URL generation
|
|
|
|
For post id `1263322`:
|
|
|
|
```python
|
|
prefix = post_id[:4]
|
|
```
|
|
|
|
Candidate originals:
|
|
|
|
```python
|
|
hosts = [
|
|
"https://rule34vault.com",
|
|
"https://r34xyz.b-cdn.net",
|
|
]
|
|
|
|
extensions = ["jpg", "png", "webp", "gif"]
|
|
|
|
for host in hosts:
|
|
for ext in extensions:
|
|
yield f"{host}/posts/{prefix}/{post_id}/{post_id}.{ext}"
|
|
```
|
|
|
|
For video direct guesses, possible candidates can be tried after HTML parsing or as fallback:
|
|
|
|
```text
|
|
<post_id>.480.mp4
|
|
<post_id>.720.mp4
|
|
<post_id>.720.hevc.mp4
|
|
<post_id>.webm
|
|
```
|
|
|
|
But actual post HTML source extraction is preferred for videos.
|
|
|
|
## CLI idea
|
|
|
|
```bash
|
|
python -m venv .venv
|
|
. .venv/bin/activate
|
|
pip install requests beautifulsoup4 tqdm
|
|
python r34vault_downloader.py krosh --start-page 1 --end-page 10 --out downloads
|
|
```
|
|
|
|
Arguments:
|
|
|
|
```text
|
|
username required
|
|
--start-page default 1
|
|
--end-page optional; stop when no items if omitted
|
|
--out default downloads
|
|
--tries default 4
|
|
--delay default 0.5 seconds between posts
|
|
--base-retry-delay default 1.5 seconds
|
|
--dry-run print planned downloads without saving media
|
|
```
|
|
|
|
## Core functions
|
|
|
|
```python
|
|
make_session()
|
|
request_with_retries(method, url, **kwargs)
|
|
fetch_user_page(username, page)
|
|
parse_ng_state(html)
|
|
extract_user_id(ng_state, username)
|
|
extract_bookmark_items_from_state(ng_state, user_id)
|
|
fetch_bookmark_items_api(user_id, page=None, skip=None, take=20)
|
|
candidate_image_urls(post_id)
|
|
stream_download(url, dest, referer)
|
|
extract_media_sources_from_post_html(post_id)
|
|
choose_video_source(sources)
|
|
already_downloaded(post_id, out_dir)
|
|
append_download_order(index, post_id, path, source_url)
|
|
append_failed(post_id, reason)
|
|
main()
|
|
```
|