Files
PFVCluster/nextprompt
T
2026-08-01 20:31:13 -05:00

105 lines
6.5 KiB
Plaintext

This is what the repo looked like before i refactored it
total 128
-rw-rw-r-- 1 reachableceo reachableceo 6311 Jul 30 23:27 AGENTS.md
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 29 19:28 awx
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 29 19:32 console
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 28 11:33 dns-cluster-setup
drwxrwxr-x 4 reachableceo reachableceo 4096 Jul 29 05:01 docs
drwxrwxr-x 4 reachableceo reachableceo 4096 Jul 29 04:59 k8s
-rw-rw-r-- 1 reachableceo reachableceo 33962 Jul 28 11:22 LICENSE
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 28 11:22 logs
drwxrwxr-x 3 reachableceo reachableceo 4096 Jul 29 19:28 netinfra
-rw-rw-r-- 1 reachableceo reachableceo 0 Aug 1 15:31 nextprompt
drwxrwxr-x 3 reachableceo reachableceo 4096 Jul 28 20:01 perf
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 28 19:46 powerman
drwxrwxr-x 10 reachableceo reachableceo 4096 Jul 30 08:47 provisioning
drwxrwxr-x 3 reachableceo reachableceo 4096 Jul 29 05:01 proxmox
-rw-rw-r-- 1 reachableceo reachableceo 2251 Jul 30 09:33 README.md
drwxrwxr-x 5 reachableceo reachableceo 4096 Jul 30 22:52 returned-logs
-rw-r--r-- 1 reachableceo reachableceo 12514 Jul 30 23:03 STATUS.md
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 30 23:01 switches
drwxrwxr-x 5 reachableceo reachableceo 4096 Jul 30 08:41 tests
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 30 11:40 ups
drwxrwxr-x 3 reachableceo reachableceo 4096 Jul 28 11:22 vendor
Here is what it looks like now
total 112
-rw-rw-r-- 1 reachableceo reachableceo 6311 Jul 30 23:27 AGENTS.md
drwxrwxr-x 4 reachableceo reachableceo 4096 Aug 1 15:34 archive
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 29 19:28 awx
drwxrwxr-x 4 reachableceo reachableceo 4096 Aug 1 15:34 dcinfra
drwxrwxr-x 4 reachableceo reachableceo 4096 Jul 29 04:59 k8s
-rw-rw-r-- 1 reachableceo reachableceo 33962 Jul 28 11:22 LICENSE
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 28 11:22 logs
drwxrwxr-x 5 reachableceo reachableceo 4096 Aug 1 15:34 netinfra
-rw-rw-r-- 1 reachableceo reachableceo 1426 Aug 1 15:31 nextprompt
drwxrwxr-x 4 reachableceo reachableceo 4096 Aug 1 15:33 proxmox
-rw-rw-r-- 1 reachableceo reachableceo 2251 Jul 30 09:33 README.md
drwxrwxr-x 5 reachableceo reachableceo 4096 Jul 30 22:52 returned-logs
-rw-r--r-- 1 reachableceo reachableceo 12514 Jul 30 23:03 STATUS.md
drwxrwxr-x 5 reachableceo reachableceo 4096 Jul 30 08:41 tests
drwxrwxr-x 2 reachableceo reachableceo 4096 Jul 30 11:40 ups
drwxrwxr-x 3 reachableceo reachableceo 4096 Jul 28 11:22 vendor
I move the top level directories into more of a sub project/module type setup
I moved the provisioning directory to archive, because its replaced by the KNELIAC repository/project (located at: /home/reachableceo/projects/KNELIAC
As far as the servers/data center itself, I have made many changes to :
- deployed virtual machines (tsys-umbrel, tsys-sipwise)
- location of virtual machine disks
- amount of ram/cpu allocated to virtual machines
- LACP between core and r3-stor is fixed
- pfv-k8s-wnode fleet changes
- removed tctc vm
- added all domains to uptime kuma
- added all systems/services to uptime kuma
- ssd/nvme disks are formatted and in use on pfv-tsys5
Unfortuantely during my storage moves, I accidently wiped out pfv-netinfra-02 . Can you turn it back into a redundant node? Its setup with the ssh pub key/passwordless sudo like before
Also, the default /dev/shm allocated to pihole was too small, i fixed that in compose live on the server, as it was a production outage. Can you fold that fix in? Also can you add a docker compose healthcheck for gravity.db issues? I had to manually fix gravitydb (move it out of way, bounce container, pihole auto recreates it). (actualy look at the bash history for root on pfv-netinfra-01 , youll see other fixes i had todo). and/or look at the docker compose config and poke around in the container. basically i had to prune volumes and bounce the container. i dont ever want to have dns fail again. this is enterprise IT , has to work.
In general, re-orient yourself to this repo, to the servers/viritual machines (via lldp/snmpwalk/ssh/nmap)
oh a note about the 4gbps lacp and concerns about mismatch.
Heres my thinking (re single flow/link saturation due to mismatched lacp and/or lack of layer4 option on core switch)....
pfv-tsys6/7 will both have 1gbps to pfv-tsys4/5 right? so thats ok
the indiviual vms (doing k8s chatter, mpi interconnect, clustered whatever) will all have 1gbps (because unique ip/mac)
So, as long as, every single endpoint has 1gbps to every other endpoint over the vlan 1000 network , I am happy.
The 4gbps was more about ensuring adequate overall capacity for all the different sessions. I am very open to more iperf testing, fio testing on the vms and generally properly performance benchmarking/analyzing the whole proxmox/k8s cluster before any workloads are put onto them. Also cpu/ram benchmarking as well. Lets see what this cluster can do.
in general, do some gardening/true up/ground truth of the servers, this repo.
utilize snmpwalk/nmap/lldp and your api access to :
- phpipam (mostly will be for the future, as i use dhcp/tailscale for everything to avoid ip conflicts)
- uptime kuma
- librenms
- beszel
- technitum
- home assistant
to get as much together as possible. let me know if anything isn't working/broke since last time. i made a bunch of changes.
I want to get you connected into promethus/promethus alert manager/grafana and have you deploy unpoller and netdisco via docker to the librenms vm (and also setup librenms with more functionality like weathermap/smokeping etc)
I want home assistant to handle all environmental monitoring and reporting for the site.
Also, any api keys must be moved into hashicorp vault as soon as possible. Track that as a very high priority todo. vault is deployed and out of box setup but nothing configured. its at https://vault.knownelement.com
Also, any work we do here, that perhaps should also be replicated in the KNELIAC repository/project, track it in a STATUS.md file in that project. We are doing a bit of fast/loose here at this time with bash scripts, in the interest of time. For the mid/long term, the work should be captured as ansible roles/playbooks.
for now, ssh to all the worker nodes (not sectestbed/preprod though) and join them to the cluster. oh the cnodes... i also accidently wiped one of them out. so that will need to be fixed. I shut the cnodes down, because one of them was 100% memory, i'm guessing cause one of the nodes got wiped out. So youll need to power them all up in proxmox.