From the original night's low-priority backlog item ("user on hold
status") — the only lever admin had for cutting a customer's access was
Revoke, which is permanent: the subscription's remaining days are just
gone, and restoring access means manually granting a brand-new one and
eyeballing how many days to give back. No way to say "block this for a
few days, then give the exact remaining time back."
db.py: new held_at column on subscriptions (same ALTER-TABLE migration
pattern as every other column added this week). hold_subscription()
sets it, guarded to only fire on a subscription that's currently active,
not already held, not expired — returns False instead of silently
no-opping so the caller can tell holding didn't happen. resume_subscription()
shifts expires_at forward by exactly how long it was held (now - held_at)
and clears held_at, so a subscription paused for 3 days comes back with
3 days added, not 3 days lost.
The part that actually mattered for correctness: list_active_subscriptions()
now also requires held_at IS NULL. This function is what xray_manager's
periodic sync (every 90s) uses to decide which clients belong in Xray's
config — without this exclusion, holding a subscription would look like
it worked for about 90 seconds and then the next sync would silently
re-add the client, since the row still has active=1 and a future
expires_at. Found this by actually tracing sync_from_db()/sync_all()
before writing the hold logic, not after debugging a live failure.
api.py: POST .../hold and .../resume routes, mirroring the existing
revoke route (fetch the sub, touch the node's xray client immediately
rather than waiting for the next periodic sync, same as revoke already
does). _days_left() now takes the whole subscription row instead of just
expires_at, so it can use held_at as the reference point instead of "now"
for a held subscription — otherwise the admin UI would show the days
counter silently ticking down while the customer isn't even able to use
the service.
admin.html: Пауза/Возобновить buttons next to Отозвать in both the main
Подписки table and the per-user card, a "на паузе" badge, and a doc-block
explaining the hold-vs-revoke distinction. Also fixed a latent race while
touching this code: the old inline revoke handler in the user card fired
openUserCard() immediately alongside revokeSub() without waiting for it,
so the card could refresh before the revoke's own API call had finished;
switched to .then() so hold/resume/revoke all correctly wait for the
action before refreshing the card.
Verification: db.py has no fastapi/aiogram dependency so this was fully
testable locally, unlike most of tonight's api.py/bot.py-touching work.
16 checks against a real isolated sqlite db: hold/resume round-trip,
the exclude-from-active-list behavior the xray sync depends on, the
exact hours-shift math (simulated a 5h hold by rewriting held_at
directly, verified the resumed expires_at landed within 6 minutes of
the expected shift), and edge cases — double-hold, double-resume,
holding an expired or already-revoked subscription, nonexistent uuid.
AST-extracted the updated _days_left() out of api.py (still can't
import the module directly) and ran it against hand-built held/active
subscription dicts. Added the same hold/resume sequence to the existing
CI "TOTP/backup/reorder" step and ran that step's exact full script
locally end to end before committing — all six of its sections pass
together, not just the new one in isolation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Neither endpoint had any brute-force protection — TOTP codes are only
6 digits (1M combinations) and HMAC-SHA1 verification is cheap, so an
unthrottled /admin/api/login/totp is a realistic brute-force target
within a pending token's 5-minute window. Password login had the same
gap.
DB-backed (new login_attempts table), not in-memory — this matters
now that mbs-api runs multiple worker processes (see 11c75c1): an
in-process counter would let an attacker split requests across
workers and bypass it entirely, same class of mistake as an
unsynchronized in-memory cache. Keyed by client IP (nginx already
sets X-Real-IP on every proxied request, install.sh has always done
this).
10 failed attempts / 15min for password, 10 / 5min for TOTP codes,
counted per-IP per-kind. Successful login clears that IP's recent
failures. Old rows pruned in the existing 90s periodic_sync cleanup
alongside sessions and pending_totp.
Verified: threshold counting, per-IP isolation, per-kind isolation
(password vs totp tracked separately), clear-on-success, and the
age-based cleanup only removing rows older than the cutoff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Remnawave's security-first positioning (passkeys/OAuth there)
with the more universal standard instead — RFC 6238 TOTP, works with
Google Authenticator/Authy/1Password/anything. Implemented from
scratch on stdlib only (hashlib/hmac/struct/base64) — no new
dependency — and verified against the official RFC 4226 HOTP test
vectors (all 10 pass exactly) before wiring it into any auth path.
Per-admin, optional: setup shows the secret + otpauth:// URI (no QR
render, just copyable text — didn't want to fake a QR library),
confirmed by entering a real code before it's persisted. Login is now
two-step when 2FA is on: password first (returns a short-lived
pending_token instead of a session if totp_secret is set), then a
second call with the pending_token + code creates the real session.
pending_totp rows are single-use and expire in 5 minutes, cleaned up
alongside the existing admin_sessions cleanup in periodic_sync.
Disabling 2FA requires re-entering the current password.
Verified end-to-end: enable/disable round trip, pending-token
resolve+single-use+expiry, code verification against the stored
secret, wrong-code and wrong-password rejection — on top of the raw
HOTP correctness check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Marzban's multi-admin (WIP there) and closes a real gap vs
both. New 'admins' table (username + PBKDF2-SHA256 password hash,
200k iterations, random salt per account, stdlib hashlib/hmac only —
no new dependency), admin_sessions now tracks which admin is logged
in. Existing installs aren't broken: on first run, if no admins exist
yet, a default 'admin' account is seeded from the current
ADMIN_PANEL_PASSWORD — old password keeps working under username
'admin', pre-filled on the login screen.
Admin management lives in Settings: list, add (username + password,
min 8 chars), remove. Can't delete the last remaining admin or your
own currently-logged-in account. Sidebar now shows who's logged in.
Verified end-to-end: bootstrap, correct/wrong/nonexistent login,
session->admin resolution, last-admin-delete protection, duplicate
username rejection, add/remove round trip, and that identical
passwords hash to different values (unique salt) but both verify.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every backup restore was leaving a .before-restore-<timestamp> safety
copy of both mbs.db and .env with no cleanup — would accumulate
forever on a panel that restores regularly. Now keeps the 5 most
recent and prunes the rest right after each restore. Verified: 8
fake copies pruned down to exactly the 5 newest, oldest-first.
admin_sessions rows also never got deleted once expired (only ever
filtered out of queries via expires_at>?) — same shape of problem.
Piggybacks on the existing 90s periodic_sync heartbeat in bot.py
instead of adding a new one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Remnawave/Marzban's 'Host Sorting via Web UI' — was the one
concrete UI gap flagged in the comparison research. Nodes get a
persistent sort_order column (backfilled once from the existing
de1-first/created_at order on migration, verified idempotent — re-running
_migrate() on an already-migrated db does not reshuffle it back).
Plain HTML5 drag-and-drop (dragstart/dragover/drop), no library —
matches the project's no-frameworks admin.html. Drop reorders the
in-memory list optimistically, re-renders immediately, then persists
via POST /admin/api/nodes/reorder; a failed save reloads from the
server instead of leaving the UI out of sync with the db.
db.reorder_nodes() rejects any list that doesn't contain exactly the
current set of node codes (no silent drops or duplicates). New nodes
append at the end (MAX(sort_order)+1) instead of jumping to the front.
Verified with a full test: initial order, reorder, two rejected
malformed reorders, a new node appending at the end, and migration
re-run not touching an already-backfilled order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
payments: _grant_paid_subscription now validates plan/node exist before
marking a payment paid instead of after (was leaving charged-but-ungranted
payments with no error trail); mark_payment_paid is now a single atomic
UPDATE ... WHERE status='pending' instead of check-then-act, closing a
double-grant race between webhooks and the periodic reconciler; yookassa
webhook now re-verifies payment status server-side via the API instead of
trusting the posted body (platega already had HMAC verification).
hwid: 'user["hwid_limit"] or FALLBACK' treated an explicit 0 (admin fully
blocking a user) as unset — now an explicit None check. Device count-check
and insert are now one atomic transaction (db.add_device_if_under_limit)
instead of two raceable statements.
perf: payment webhooks and _grant_paid_subscription's SSH/HTTP calls now
run via asyncio.to_thread instead of blocking the event loop; same for
bot.py's periodic_sync/reconcile_pending_payments and the manual admin
sync button. Admin endpoints (traffic/subscriptions/payments/gift-codes/
user-card) now resolve node labels from one db.list_nodes() call instead
of a fresh db.get_node() per row. revoke/reset-traffic use a direct PK
lookup instead of scanning up to 5000 rows. Dashboard now asks the API
for 8 rows instead of fetching 200 and slicing client-side.
security: mbs.db (and -wal/-shm) now chmod 600 right after creation —
it held session tokens and subscription bearer tokens world-readable
by default. Node SSH connections now pin host keys via a persisted
known_hosts file (TOFU) instead of accepting any key on every connection.
delete_node now refuses to delete a node with active subscriptions
instead of silently orphaning their xray clients.
deadcode: removed unused xray_manager.list_client_ids and admin.html's
superseded staggerReveal (rows animate via rowAttr() inline now).
Also guards gift-code redemption against a plan/node deleted after the
code was created (was an unhandled KeyError/TypeError crash).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>