Lattice NPM
Training & introduction
Intro for a NOC operator, engineer, or executive who has never opened Lattice. Start with the 15-minute path. Then run the two playbooks.
First 15 minutes
Do this once. You will know if you can page, and you will have touched the four screens that matter.
- Overview — look at the strip. Red is Critical. Yellow is Degraded. If both are empty, the fabric is quiet. Site tiles underneath; click HQ.
- Map — Internet at the top, firewall, then switches, hosts, VMs. Hover HQ-ACC-C9300. That card is the inventory. Double-click opens the dashboard.
- One device — HQ-FW-PA3220. Health ring, DIA-facing interfaces, setup pills (ICMP/SNMP). If the ring is green and DIA is yellow, the circuit is the story, not the box.
- One report — Reports → Element is already on DIA. Read the headline paragraph. That is the report. Saved views: DIA month, Site month, Worst 10.
- Alerting — Policies tab is the six rules we page on. Ack a minor. Outbox is the demo email/SMS. Nothing leaves the building.
1. What Lattice is
Lattice is our Network Performance Monitor. It is a new system — not The Dude, not SolarWinds, not PRTG. It does the same job: tell you whether the network is healthy, where it is not, and why, without SSH into fifteen boxes.
- Watches nodes, circuits, Hyper-V hosts/guests, APs, and apps we have loaded.
- Rolls that into a health score (0–100) so you are not staring at raw SNMP.
- Drills fabric → site → neighbor → node.
- Reconstructs the path a user takes (NetPath) and overlays metrics (PerfStack).
- Stores downtime, SLA, and outage patterns — Tuesday 02:00 is a pattern, not a surprise.
- Imports inventory, including a Dude backup, so we do not re-type the network.
2. Who uses which part
| Role | You live here | You rarely need |
|---|---|---|
| NOC | Overview, Map, NOC, Alerting, Events | Config diffs, UnDP |
| Engineer | NetPath, PerfStack, Insight, Configs, Hyper-V, Reports → Element | Executive rollup |
| Admin | Admin (nodes, links, Dude import), Teams | QoE |
| Exec | Overview health, Reports → Executive, Capacity | Topology drill |
SSO is Entra ID (demo-connected). Roles: Admin, NOC, Engineer, Exec, Read-only. This preview is not locked behind login. On the live server, SSO gates the same screens.
3. The language of the board
Learn this before you click anything. Same colors on map, alerts, reports, Hyper-V.
| You see | Means | Do this |
|---|---|---|
| Green / Up | Responding. Health typically ≥ 90. | Nothing. |
| Yellow / Degraded | Up, but not well (CPU, loss, pressure, crowded AP, cluster n-1). | Investigate before it is an outage. |
| Red / Down / Critical | Not responding, circuit down, or Critical alert. Health 0. | Incident. |
| Grey / unknown | Not polled, or VM Saved/Off. | Confirm it should be monitored. |
Header, every page: live health, up / degraded / down. If health is not in the 90s, start at Overview or Alerting.
4. Health score
A node is not “up” just because ping works. Five slices, then a rollup:
- Availability — ICMP/SNMP. Zero if down, lost heartbeat, or VM Off/Saved.
- Performance — CPU, memory, Hyper-V CPU-wait, balloon.
- Hardware — temp, fans, PSUs; VHDX latency and host free RAM on Hyper-V.
- Response — ICMP vs that node’s own baseline (1→8 ms is worse than a branch that always sits at 30 ms).
- Interfaces — oper-down, high util, errors.
90–100 healthy · 70–89 degraded · below 70 or 0 treat as down. A SQL VM can be green on ICMP and yellow on Apps. Check the layer you care about: network vs hypervisor vs application.
5. The demo fabric
Unless someone imported a Dude file, you are on Northwind:
- HQ Chicago — Palo Alto edge, Cisco access, Arista leaf, two 9800 WLCs, NAS, 2-node Hyper-V (HQ-HV-01 / 02) with CSV, SQL, DC, RDS, PBX.
- Austin / Denver — Cisco ISR over IPsec. APs join HQ-WLC-9800-1 (FlexConnect).
- DIA — two circuits: ISP-PE-CHI (DIA-1 2G) and ISP-PE-CHI-ZAYO (DIA-2 1G). Report DIA-1 first.
- Wi-Fi world — NYC, São Paulo, London, Frankfurt, Dubai, Singapore, Tokyo, Sydney. Each has a site ISR and 2 ISPs (NYC / London / Singapore / Denver have a 3rd). APs CAPWAP to the two controllers over that local DIA.
- M365 as a QoE target, not an SNMP box.
Inject from Overview or Hyper-V: WAN loss, AP down, host contention, cluster failover. They expire. That is the lab.
6. First ten minutes
Same path as the 15-minute card, written as a checklist. If you already did that, skip this.
- Overview — Critical/Degraded strip, site health, DIA, capacity days-to-exhaust.
- Click a site tile — Map opens already drilled.
- Map: Internet → firewall → switch → host/VM. Hover for inventory. Double-click dashboard.
- Open HQ-FW from Devices. Setup pills, probes, interfaces.
- NetPath — preset DIA (HQ-FW → ISP-PE). First red hop is where the user feels it.
- PerfStack — pick DIA, Core, SQL, or HV cluster. Do not build the first overlay.
- Hyper-V — two hosts, CSV owner, n-1 banner if a node is down.
- Reports → Element. It opens on DIA. Saved: DIA month, Site month, Worst 10.
- Admin — look, don’t rewrite production IPs. Reset demo restores Northwind.
- Alerting — Policies (six defaults), ack a minor, read the outbox.
7. Monitor
Overview
Operations home. Critical/Degraded strip first, then health, DIA, capacity, sites. Shortcuts sit under More. Run discovery is simulated. Inventory work lives in Admin.
Map
Tree from Internet down. Hover a node or a wire. Time travel plays snapshots back to the break.
NOC board
Six tiles: Sites, Open incidents, DIA, Cluster, Worst node, Alert age. Troubleshoot from Map / NetPath / the node, not from the TV.
Devices
Inventory columns: Status, Name, IP, Kind, Site, Health, Uptime. Node page has setup pills (ICMP/SNMP/WMI).
NetPath
Default path is DIA. Presets for Austin and Denver. Healthy firewall ping plus a red hop upstream is underlay/ISP, not LAN.
PerfStack
Presets: DIA, Core, SQL, HV cluster. Two lines spiking together are related until proven otherwise.
Wireless
APs, clients, noise. HQ-AP-FL1 at 27 stations vs baseline ~16 is a real seed example. Warehouse scanners stay on their VLAN.
Hyper-V
Two hosts, one cluster, CSV owner, witness, HA vs parked guests. Host: logical CPU, free RAM, VHDX, balloon. Guest: vCPU wait, assigned vs demand, heartbeat, state. Parked VMs must show “—”. When HV-01 is down, HV-02 yellow is correct (cluster n-1), not a second incident. SQL/DC/RDS also appear under Apps.
Experience
Off the sidebar. Application share and RTT still exist at /qoe if you need them. Pair with Traffic when “the internet is slow.”
8. Analyze
Traffic
NetFlow top talkers: who filled DIA, who is pounding CSV. Unknown UDP at the firewall is worth a look.
Configs
Running vs startup diff, CIS-style compliance. Austin shop NAT drift is the teaching example. Come here after a change window before you reload the box.
Apps
SQL (batch, page life, locks, log flush), AD (LDAP bind, replication), RDS (sessions, input delay). Page life falling while the VM looks fine is still a DBA incident — and you can prove the host was not starving it.
IPAM & UDT
Subnets, used/reserved/conflict IPs, and switch-port → MAC → user. “Who is on Gi1/0/8?” lives here. Demo includes a printer/laptop conflict on purpose.
Insight
PAN sessions, DP CPU, IKE tunnels with RTT, drop reasons, BGP. ICMP to the firewall is not “IKE is up.”
Capacity
Exhaustion dates. SQL RAM ~18 days should scare you; DIA ~47 days is the ISP conversation. Baselines: high-but-normal for that box is not an incident.
Events
Stream (syslog/traps) · Maintenance (toggle windows) · Dependencies (parent→child suppress) · Pollers + custom OIDs.
9. Operate
Alerting
Critical red, degraded yellow. Channels write to the outbox in demo — they do not leave the network. Ack clears the badge; history stays.
Reports
Element is the default. Filter DIA / routers / switches / … Read the headline paragraph first — that is why you ran it. Then current KPIs, history, this object’s outages only. Start DIA. Click a worst-node on Downtime to jump into its Element report.
Other tabs: Downtime (SLA, MTTR, MTBF), Patterns (heatmap, Tuesday 02:00, flapping, blast radius), Availability vs 99.95%, Executive (Monday email button), Capacity/Traffic.
Teams
Directory, role matrix, SSO connected/paused. Preview stays open.
Admin
Nodes (form editor — table is read-only so polling doesn’t steal focus), Links, Import. Merge is default (keeps Hyper-V). Replace wipes the fabric — training should not. Reset demo restores Northwind.
10. Playbook: working an incident
- Header — health not in the 90s? Down vs degraded.
- Alerting — newest Critical. Read it. Do not ack yet.
- Map, fabric — one site or the WAN? DIA red → don’t chase Austin.
- Events → Dependencies — suppressed? Work the parent.
- NetPath — first red hop is the fault domain.
- PerfStack — backup window vs ISP vs host pressure.
- VM? Hyper-V host wait / CSV / heartbeat, then Apps.
- WAN? Insight (IKE, BGP) + Traffic + Reports Element on DIA.
- Configs — commit in the last hour?
- Ack. Write the cause in manager language: “DIA loss 1.8%. Austin is downstream, not a branch fault.”
11. Playbook: running a report
- Reports → Element.
- Chip DIA (or the ISR they named).
- Read the headline. If it answers the question, paste it into the ticket and stop.
- 30-day availability and this object’s outages — not the fabric log.
- “Every Tuesday?” → Patterns heatmap / recurring windows.
- SLA language → Downtime or Availability vs 99.95%.
- Email downtime report is the exec brief (demo outbox).
More data is there if you scroll. Do not send a 40-row table when the paragraph already said “two ISP events, 41 minutes, SLA breach.”
12. Importing from The Dude
Lattice is not Dude. Dude is an inventory source.
- Admin → Import.
- Leave mode on Merge unless you were told to replace.
- Drop the backup (XML / export / sqlite) or Load Dude sample.
- Read the toast: counts and warnings.
- Spot-check Map, a known IP, and that the Hyper-V cluster is still two hosts if you merged onto Northwind.
Ugly names (`device-12`) get renamed in Admin → Nodes. Lattice keeps the IP.
13. Calibrate (color → cause)
| Symptom | Color | Where | Typical cause |
|---|---|---|---|
| DIA util 50s, loss 0.2% | Green/yellow | Reports DIA | Backup window, not down |
| DIA loss > 1% | Red alert | Alerting + Insight | ISP / WAN inject |
| HV-01 wait, guests balloon | Yellow | Hyper-V | Host contention inject |
| HV-01 down, SQL migrating | HV-01 red, HV-02 yellow | Hyper-V | Failover — one incident |
| AP 27+ clients | Yellow | Wireless | Overcrowd inject |
| Austin down, DIA down | Austin suppressed | Events deps | Work DIA |
| SQL page life low, host fine | Apps yellow | Apps | Database, not compute |
| Saved VM | Grey, metrics — | Hyper-V | Parked — do not “fix RAM” |
14. Demo vs live
| Capability | Here | Live poller |
|---|---|---|
| Screens, health, maps, reports | Product | Same |
| SNMP / ICMP / WinRM | Simulated tick | Real poll |
| Traps / syslog / NetFlow | Simulated | Real receivers |
| Email / SMS | Outbox preview | SMTP / SMS gateway |
| SSO | Directory, no login wall | Entra SAML gate |
| Dude import | Parser + merge | Same parser |
| Inject buttons | Training | Disable or keep staged |
If a number looks too perfect, it is the simulator. You are learning the workflow.
15. Glossary
| Term | Meaning |
|---|---|
| DIA | Dedicated Internet Access — HQ’s circuit to the ISP. |
| Health | 0–100 composite, not ping. |
| NetPath | Hop-by-hop path and quality. |
| PerfStack | Multi-metric overlay on one time axis. |
| CSV | Cluster Shared Volume (Hyper-V). |
| CPU wait | Guest ready but waiting on host CPU. |
| Balloon | Host reclaiming guest RAM. |
| UDT | Port / MAC / user. |
| UnDP | Extra OID poller. |
| MTTR / MTBF | Mean time to restore / between failures. |
| SLA | 99.95% unless leadership says otherwise. |
| Merge vs Replace | Import additively vs wipe the fabric. |
| Suppressed | Muted because parent is down or maintenance is on. |