docs: archive superseded question rounds to archive/questions [#769]

v10 is the current round. Old rounds preserved in git history and
archive/questions/.
This commit is contained in:
2026-09-04 05:57:04 -05:00
parent 823eb9b33d
commit ce55409b9a
6 changed files with 0 additions and 0 deletions
+135
View File
@@ -0,0 +1,135 @@
> Working ask-file; knowledge docs live on Discourse: https://community.turnsys.com/c/vp-techops (audit log: https://community.turnsys.com/t/298)
# questions-09022206.md — consolidated open asks (2026-09-02)
> Merged from `questions-v2.md` (repo root) + `HomeAssistant/needfromcharles.md`
> per founder request; both source files since removed (history in git).
> Answer inline; agent checks before asking again.
## HA / monitoring (was needfromcharles)
### 1. Minisplit (Midea/Carrier) — blocks #614
Login `charles@turnsys.com` rejected by all three Midea cloud backends.
- [ ] Exact app name on your phone that controls the minisplit:
- [ ] Account email that app uses:
### 2. Govee — blocked on hardware path (#620/#683)
Integration live (key+account+MQTT); account has ZERO cloud devices = BLE-only.
- [ ] Sensor model numbers (decides BLE dongle vs hub):
- [x] API key — received + configured (2026-09-02)
### 3. HomeKit codes — blocks #619 pairing (ecobee multicast-reachable NOW)
- [ ] ecobee 8-digit code:
- [ ] Lock brand/model: (verdict 2026-09-02: BLE-only — DHCP can't name it)
- [ ] Lock HomeKit code:
### 4. Emporia Vue 3 — blocks #623 (BTU loop)
- [ ] Ordered? ETA:
### 9. iDRAC — RESOLVED 2026-09-02
[x] Creds worked; both SNMP agents live; tsys6 `public` accepted by founder.
### NEW (2026-09-02)
### 10. DVR → MotionEye (#711)
- [ ] DVR login (any user):
- [ ] Channel count:
(Night Owl; RTSP format documented on #711)
### 11. DIRIGERA Matter code (#712) — when you're back at the house
- [ ] IKEA app → hub → integrations → share via Matter (11-digit code)
### 12. WiZ bulb→room mapping (#712) — Friday onsite
2 bulbs reserved (pfv-wiz-bulb-1/2 @ 192.168.1.150/.151); map MACs→rooms, then
HA WiZ entries by host + we identify the other ~11 lights' integration.
## General (was questions-v2)
### Q7. Close #167 (old garage UPS rebuild, superseded)? — still open
### Q9. Session-start gate relaxation — superseded by scope-alignment gate
### Q10. Git flow: keep direct-push for PFVCluster? — still open (default: keep)
(Q8 Friday order — moot, that Friday passed)
## Founder decisions pending
- #626 Cloudron HA access architecture (a/b/c on ticket — recommend proxy)
- #684 PBS session scheduling (+ drive-rotation target design)
- TasmoAdmin uninstall (stopped+disabled 2026-09-02 — full uninstall?)
- Govee disabled dup entry: one-click delete in HA UI
### 13. Cloudron API token (#685/#727) — unblocks API-first access
The #685 perf investigation ran over raw SSH (root@my.knownelement.com) because
no panel API token exists. A token (read scope minimum) lets the agent use the
Cloudron API (apps/configs/backups) instead of shell — cleaner audit trail and
the path for AWX/Ansible modules. Also tracked in KNEL/cloudron
questions-v1.md (https://git.knownelement.com/KNEL/cloudron).
- [x] Create token (panel: my.knownelement.com → profile → API access) — received 2026-09-02
- [x] Material value stored → ~/.creds/cloudron.env (0600; API verified: profile mrcharles, apps 200)
### 14. docker daemon.json bounce window — my.knownelement.com (#727)
One planned dockerd restart (ALL ~300 containers bounce; live-restore=false
today, so no zero-downtime path) to apply nofile 262144 defaults + log
rotation + live-restore. Script ready: KNEL/cloudron
perf/docker-daemon.json.planned. Until then the postgres nofile bridge is
ephemeral (lost on postgres container recreate). Detail: t/316.
- [x] Pick a window — DONE 2026-09-02 ~19:30 CDT: applied (live-restore + nofile 262144 defaults); first-attempt log-driver conflict incident documented on #731
## CMDB change-control rulings needed (#705 P0 live → P2/P3)
### C1. CR itemtype for the 7-state workflow
GLPI natively has both Ticket and Change. The t/319 workflow
(Draft→Review→Approved→Scheduled→Implemented→Verified→Closed) maps most
cleanly onto GLPI **Changes** (validation/plan/approval fields built in),
with Tickets for break-fix. Recommend: Changes for CRs, Tickets stay for
incidents; both linked to the CI they touch.
### C2. Write access for the cmdb agent in the CR phase
Agent is Read-Only today (verified). If CRs live in GLPI, the agent needs
to CREATE/transition them — recommend adding the Self-Service profile
alongside Read-Only (can open + follow own items, still can't edit CIs).
Alternative: CRs enter via Redmine only and GLPI stays read-only for now.
### C3. #307 → GLPI refresh cadence
from-inventory.sh regenerates the CSV from the Discourse #307 raw;
to-glpi.sh is idempotent by name. Recommend folding "refresh #307
regenerate → import → report diff" into the daily 09:00 C2 housekeeping.
Related data-quality: #307 §5 has `pfvsvrki.knel.net` typo; 14
stale-dns-record CIs need reconcile-or-delete ruling.
## Agent-stack rulings needed (#767 — after TSGCOO tree discovery 2026-09-03)
### C4. Credentials policy — one statement, please
Three coexisting models: Bitwarden (TSGCOO's bw-run.sh design — replaces ~/.creds for
agent/system creds), ~/.creds 0600 (current house rules), Vaultwarden (your app-secrets
ruling). Proposed: Bitwarden = agent + system API credentials going forward (~/.creds
deprecated as #478 ingests them), Vaultwarden stays for app-internal secrets. Confirm?
> **ANSWERED 2026-09-03 (Charles):** Bitwarden/Vault is the SoR — "its the latter."
> ~/.creds is a bootstrap mechanism for convenience only (same as using the
> reachableceo linux/cloudron account as a bootstrap identity). Going forward: new
> agent/system credentials land in Bitwarden; ~/.creds materializes bootstrap copies.
### C5. #767 scope split
This chat owns the CR escalation ladder + GLPI Changes wiring + mechanical gates
(ticket-gate CR check, X-Consumer attribution); TSGCOO's lane owns identity provisioning
execution (#442/#478). Confirm the seam?
> **ANSWERED 2026-09-03 (Charles):** No split — **this chat is COO/C2 only going
> forward** (project management, policy making, enablement). The agent stack IS
> enablement → it lives HERE, end-to-end (design + provisioning execution). Dedicated
> chats get spun up for any work that isn't PM/policy/enablement. TSGCOO's existing
> tree is prior art feeding this chat, not a separate lane.
### C6. Resume TSGCOO's blockers?
Q1-Q6 in org-buildout/questions-v1.md + #478 stage-2 re-gate. I can clear Q1 (docker
group for TSGCOO) now — say the word. Invite links (Q3) + Gitea token (Q4) are yours.
> **ANSWERED 2026-09-03 (Charles):** "yes give all of the TSG* and COS* accounts docker
> group access. that is super critical for all them to work." → VERIFIED ALREADY IN
> PLACE: docker(988) confirmed on all 7 officer accounts (TSGCOO, TSGCCO, TSGCTO,
> TSGBOD, COS-RCEO, COS-WFO, COS-TSG). Org-buildout Q1 closed. Q3 (Cloudron invite
> links ×9) + Q4 (Gitea push token) remain Charles's.
### C7. Leaked vptechops Gitea token
Listed in a KNELSecretsManager remote URL — rotate at your convenience (your console,
not my lane).
### C8. GLPI Changes vs Tickets (was C1)
Still the blocking ruling for the whole CR wiring — see C1 above (recommend Changes).
+55
View File
@@ -0,0 +1,55 @@
# Questions — v4 (current round)
v3 closed 2026-09-02: Q1 answered in chat ("cloudron postgres solved —
wire up the remote access, hook first"). The deploy hook LANDED (#734,
E2E proven); the pairing hit a wrong premise and needs a fork ruling
below. Version up (v4 -> v5) on the next Q&A round.
## Active question channel (founder-directed)
`~/PTKR.md` — 12 open items: GLPI API key, AI-staff identities,
Vault-vs-Bitwarden for agent secrets, prometheus/grafana envs, VoIP carrier
decision, PMG MX cut-over window, SITES access, October hardware window
dates, Friday 2026-09-04 onsite coordination, #681 Cloudron 500s, nitrokey
purchase, Cloudron packaging session scope.
## Repo-local questions
### Q1 — HA instance pairing mechanism (blocks #626 family access; would like before Friday)
The 2025.7+ "built-in pairing" premise was WRONG — no such core
integration exists as of 2026.8.3 (verified on both instances + HA core +
docs; full research in #626 note). Options:
- **(A)** HACS custom `remote_homeassistant` on the Cloudron HA — full
entity mirroring + control, matches the ruled architecture. Needs a
founder file/HACS install on the Cloudron app (~30 min of your time;
no agent API for Cloudron app files). Caveat: community reports of
recent breakage (remote entities unavailable) — I'd vet the GitHub
issues before committing.
- **(B)** One-way API bridge — pfv-bms pushes ~20 key family states
(temps, UPS, garage/artroom) to the Cloudron HA REST API + I build
family dashboards there via websocket. I can build it entirely
tonight; family gets read-only visibility; no control.
- **(C)** MQTT broker on the VPS, both instances bridge — bidirectional,
but a new always-on service + YAML on both sides.
Answer (A/B/C) + does family need to TOGGLE things at the property, or
just see status?
### Q2 — gitea webhook allowlist (1-line + restart; unlocks instant deploys)
gitea silently drops webhook deliveries to CGNAT-range targets (SSRF
`ALLOWED_HOST_LIST` default — no delivery attempt, no log line; restart
doesn't help; that's why hook #10 never fired). Until it opens, the
5-minute sha-watch net deploys everything reload-only — already live.
To open the instant path: Cloudron file manager → gitea app.ini →
```ini
[webhook]
ALLOWED_HOST_LIST = external,private
```
then restart gitea. Say the word when done — I'll verify hook #10 fires
end-to-end (it carries the full restart-mapping logic).
+35
View File
@@ -0,0 +1,35 @@
# Questions — v6 (current round)
Round v5 resolved 2026-09-02 by founder chat ruling + agent action;
carried below. One question remains open (Q4).
## Resolved
**Q1. Door actuator path?**
**A:** Same exact hardware from 2018 — usbrelay strike on pfvsvrpi.
Relay dongle confirmed present (16c0:05df); usbrelay tool install is a
ticket-gated #356 item.
**Q2. Reader model + badge format?**
**A:** Probed 2026-09-02 (founder cleared): both readers = 13ba:0018
PCPlay Barcode PCP-BCG4209; pfvsvrpi aarch64/event0, ultix-field
x86_64/event15; 24-byte records; 10-digit decimal + Enter format.
**Q3. Badge whitelist home?**
**A:** HA-side whitelist, confirmed. Now materialized as a separate
repo (see Q5).
**Q5. Same 2-approval gate on the badge roster home (pfv-bms)?**
**A:** RESOLVED 2026-09-02 — founder chat ruling: the human gate is for
badge IDs only, and the roster should be 100% separate. Built:
**KNEL/access-roster** (private) — PR-only + two human approvals,
enforced; holds `roster/badges.yaml` (the HA control file). doorman
code repo de-gated for fast movement. HA consumption wiring owned by
the pfv-bms session.
## Open
**Q4. Second human approver in gitea?**
**A:** Founder: "planned for very near future." Until granted + named,
KNEL/access-roster merges are IMPOSSIBLE by design (fail-safe). Does
not block code/deploy work.
+76
View File
@@ -0,0 +1,76 @@
# Questions — v7 (current round)
Round v6 carried below (badge Q4 still open). New this round: the
Proxmox Mail Gateway pair (#696, child of #694) design questions
PQ1PQ7. Recommended defaults are stated per question — answering
"defaults fine" (or naming exceptions) is enough.
## PMG round (#696) — awaiting answers
**PQ1. Mail domains + backend target?**
Which domains' mail should flow through the PFV PMG pair (e.g.
knownelement.com, turnsys.com, thefnf.net)? And what sits behind PMG
as the delivery target — the Cloudron VPS in Reston VA? If so, which
hostname should PMG smarthost to (e.g. `mail.<domain>`)?
*Ground truth: the internal Technitium `knel.net` zone has no MX
records; public MX lives at OVH/Cloudron (founder DNS).*
**PQ2. Inbound path from the internet?**
Recommendation: edge DNAT of SMTP :25 → both PMG nodes (LAN IPs),
PMG filters then relays to the Cloudron backend. Which edge does the
DNAT live on today — the rrinfra Cisco (existing), and are there any
existing mail port-forwards I should be aware of?
**PQ3. Outbound smarthost?**
Should LAN/Tailscale systems (hosts, VMs, appliances) relay outbound
mail through PMG? If yes, confirm authorized source subnets for the
allowlist. And Cloudron's own outbound: keep it direct (as today) or
route it back through PMG too?
**PQ4. VM placement + sizing?**
Recommendation: 2 VMs, each 2 vCPU / 4G RAM / 32G system disk (PMG is
light). PMG-01 on tsys4, PMG-02 on tsys6 — honoring the standing
ruling that tsys6/7 use USB storage only (PMG disks are small, fits
fine). VMIDs: next free in the production range unless you want
specifics. OK?
**PQ5. Native PMG cluster or independent pair?**
Recommendation: PMG native clustering (config + quarantine DB
replicates between the two), managed strictly serial with a health
gate between hops per the blue/green house rule. OK?
**PQ6. TLS from the fleet CA (#697)?**
Recommendation: issue certs for `pmg-01.knel.net` / `pmg-02.knel.net`
from the tsys-ca intermediate for the admin UIs; SMTP keeps
opportunistic TLS (external senders never see our internal CA, so no
interop risk). OK?
**PQ7. Quarantine + admin notifications?**
Defaults I'd apply unless you say otherwise: quarantine retention 7
days, daily quarantine digest to your admin address — **which address
should receive quarantine reports and PMG alerts?**
## Carried from v6 — still open
**Q4. Second human approver in gitea?**
**A (from v6):** Founder: "planned for very near future." Until
granted + named, KNEL/access-roster merges are IMPOSSIBLE by design
(fail-safe). Does not block code/deploy work.
## Carried from v6 — resolved
**Q1. Door actuator path?** Same 2018 hardware — usbrelay strike on
pfvsvrpi (16c0:05df); usbrelay tool install is ticket-gated #356.
**Q2. Reader model + badge format?** Both readers = 13ba:0018 PCPlay
PCP-BCG4209; pfvsvrpi aarch64/event0, ultix-field x86_64/event15;
24-byte records; 10-digit decimal + Enter.
**Q3. Badge whitelist home?** HA-side whitelist, materialized as
separate repo KNEL/access-roster (see Q5).
**Q5. Two-approval gate on roster home?** RESOLVED 2026-09-02 —
founder ruling: roster is 100% separate. Built: KNEL/access-roster
(private, PR-only + two human approvals); holds roster/badges.yaml.
doorman code repo de-gated. HA consumption owned by the pfv-bms
session.
+69
View File
@@ -0,0 +1,69 @@
# Questions — v8 (current round)
v7's PMG round is partially resolved — the 2026-09-03 mail audit (55
systems, read-only) answered several questions with ground truth and
Charles answered others in chat. Answers recorded below; new questions
PQ8PQ10 need your go. Full map + deploy plan:
[Discourse t/322](https://community.turnsys.com/t/322).
## PMG round — resolved
**PQ1 (backend).** ✅ Cloudron VPS confirmed as backend:
`mail.knownelement.com` (152.53.37.179). Audit adds: public MX already
points there; fleet delivers direct today; knel.net zone has no MX.
**PQ3 (outbound smarthost).** ✅ Charles ruled: all hosts send through
a shared **VIP** in front of the PMG pair. (Mechanics in PQ8 below.)
**PQ4 (placement/sizing).** ✅ Resolved by reality: the pair already
exists — VM 604 on tsys6 (192.168.1.11) and VM 711 on tsys7
(192.168.1.10), both 2vCPU/4G/32G, running, onboot, PMG 9.1.0
installed, unclustered, default config. Nothing to build; configure
only.
**"Old relay in postfix?"** ✅ None. `relayhost` is empty on every
audited system (16 postfix installs); no ssmtp/msmtp/nullmailer
anywhere. Fleet sends direct-to-MX today.
## PMG round — still open
**PQ2 (inbound path).** Where does public :25 DNAT land — rrinfra
Cisco? Any existing mail port-forwards? (Founder-side; only needed
before the MX cut-over, not for outbound value.)
**PQ5 (native PMG cluster?).** Recommended: yes — config + quarantine
replication between .1.11/.1.10, managed strictly serial. OK?
**PQ6 (TLS from fleet CA #697?).** Recommended: yes for the two admin
UIs; SMTP stays opportunistic. OK?
**PQ7 (notify address).** Evidence says the pair already forwards to
`coo@turnsys.com`. Confirm that address for quarantine digests + PMG
alerts?
**PQ8 (VIP specifics).** Proposal: **VIP 192.168.3.249** (pool ends
.3.200; .250-.254 is your infra block), keepalived VRRP VRID 10 on the
pair, dhcpd host-declaration pin (VRRP MAC 00:00:5e:00:01:0a), DNS A
record `smtp.knel.net` → 192.168.3.249. Confirm the number + name, or
name your preferred slot.
**PQ9 (outbound cutover staging).** Recommended: serial staged — PVE
hosts + PBS first (verify a week), then the VM postfix set, one
blue/green hop at a time. OK, or all at once?
**PQ10 (side findings — sub-tickets?).** (a) tsys-ucs-02 root alias
blackholes to `systemmail@unassigned-hostname...`; (b) Wazuh alerting
never configured (stock example smtp). Fix as ride-alongs in #696 or
separate sub-tickets? (Will create them either way — just tell me
where they land.)
## Carried from v6 — still open
**Q4. Second human approver in gitea?** "Planned for very near
future"; until then KNEL/access-roster merges are impossible by
design. Does not block code/deploy work.
## Carried from v7 — superseded questions
PQ1/PQ3/PQ4 answers above replace the v7 wording; v7 file remains in
git history.
+69
View File
@@ -0,0 +1,69 @@
# Questions — v9 (current round)
v8's open items are now settled by the founder's 2026-09-03 chat
ruling: **OUTBOUND ONLY** — all PFV hosts (k8s included) send via the
gateway; no inbound ever; no DNAT, no MX cut-over (that scope is
dropped from #696). Robustness is a stated requirement (VIP pair,
queueing, cluster, monitoring). Map + target architecture updated in
place: [Discourse t/322](https://community.turnsys.com/t/322).
## PMG round — resolved
**PQ1 (backend).** Cloudron VPS `mail.knownelement.com`
(152.53.37.179); public MX already points there; fleet delivers
direct today; knel.net zone has no MX (and never needs one now).
**PQ2 (inbound path).** ✅ N/A by ruling — outbound only. No DNAT, no
MX work, ever.
**PQ3 (outbound smarthost).** All hosts through the gateway VIP.
**PQ4 (placement/sizing).** Pair already exists: VM 604 tsys6
(192.168.1.11) + VM 711 tsys7 (192.168.1.10), PMG 9.1.0, unclustered.
**"Old relay in postfix?"** None — `relayhost` empty fleet-wide.
## PMG round — still open (one-line answers fine)
**PQ5 (native PMG cluster?).** Recommended: yes — config + quarantine
replication between .1.11/.1.10, managed strictly serial. OK?
**PQ6 (TLS from fleet CA #697?).** Recommended: yes for the two admin
UIs; SMTP stays opportunistic (all flows are internal + to Cloudron).
**PQ7 (notify address).** Pair already forwards to `coo@turnsys.com`.
Confirm for PMG alerts + any held-mail notices?
**PQ8 (VIP specifics).** Confirm **VIP 192.168.3.249** + name
`smtp.knel.net` + dhcpd host-decl pin (VRRP MAC 00:00:5e:00:01:0a)?
Or name your preferred slot/name.
**PQ9 (outbound cutover staging).** Recommended: serial staged — PVE
hosts + PBS first (verify a week), then VM postfix set. OK?
**PQ11 (Cloudron relay user — founder side).** Recommended pattern:
PMG authenticates to a dedicated Cloudron mail user (e.g.
`pfv-relay@turnsys.com`) over :587 submission, so fleet mail exits
with Cloudron's rDNS/reputation. Will you create that user (or point
me at who does)? Fallback if you'd rather not: PMG delivers direct to
recipient MXs (same egress as today, no deliverability gain).
**PQ12 (k8s submission port).** Recommended: :25 unauthenticated from
trusted LAN (pod traffic SNATs to node IPs) as the default; optional
:587+auth later if you want app-level credentials. OK as default?
**PQ10 (side findings — sub-ticket placement).** (a) tsys-ucs-02 root
alias blackholes to `systemmail@unassigned-hostname...`; (b) Wazuh
alerting never configured (stock example smtp). Ride-alongs in #696
or separate sub-tickets? They get tickets either way — just say where.
## Carried from v6 — still open
**Q4. Second human approver in gitea?** "Planned for very near
future"; until then KNEL/access-roster merges are impossible by
design. Does not block code/deploy work.
## Superseded
v7 PQ wording and v8 PQ10 ordering are superseded by the rulings
above; v7/v8 files remain in git history.