Lattice
Health 9849 up1 down

Lattice NPM

Training & introduction

Intro for a NOC operator, engineer, or executive who has never opened Lattice. Start with the 15-minute path. Then run the two playbooks.

First 15 minutes

Do this once. You will know if you can page, and you will have touched the four screens that matter.

  1. Overview — look at the strip. Red is Critical. Yellow is Degraded. If both are empty, the fabric is quiet. Site tiles underneath; click HQ.
  2. Map — Internet at the top, firewall, then switches, hosts, VMs. Hover HQ-ACC-C9300. That card is the inventory. Double-click opens the dashboard.
  3. One device — HQ-FW-PA3220. Health ring, DIA-facing interfaces, setup pills (ICMP/SNMP). If the ring is green and DIA is yellow, the circuit is the story, not the box.
  4. One report — Reports → Element is already on DIA. Read the headline paragraph. That is the report. Saved views: DIA month, Site month, Worst 10.
  5. Alerting — Policies tab is the six rules we page on. Ack a minor. Outbox is the demo email/SMS. Nothing leaves the building.
Three clicks if you remember nothing else: Overview strip → Map → Reports Element (DIA). PerfStack presets (DIA / Core / SQL / HV cluster) if “it is slow.”

1. What Lattice is

Lattice is our Network Performance Monitor. It is a new system — not The Dude, not SolarWinds, not PRTG. It does the same job: tell you whether the network is healthy, where it is not, and why, without SSH into fifteen boxes.

  • Watches nodes, circuits, Hyper-V hosts/guests, APs, and apps we have loaded.
  • Rolls that into a health score (0–100) so you are not staring at raw SNMP.
  • Drills fabric → site → neighbor → node.
  • Reconstructs the path a user takes (NetPath) and overlays metrics (PerfStack).
  • Stores downtime, SLA, and outage patterns — Tuesday 02:00 is a pattern, not a surprise.
  • Imports inventory, including a Dude backup, so we do not re-type the network.
This environment is a high-fidelity demo fabric. Screens are the product. Numbers tick from a simulator. On a live poller the same screens fill from SNMP, ICMP, WinRM, traps, and NetFlow. Voice / IP SLA is built but turned off — it is not in the sidebar.

2. Who uses which part

RoleYou live hereYou rarely need
NOCOverview, Map, NOC, Alerting, EventsConfig diffs, UnDP
EngineerNetPath, PerfStack, Insight, Configs, Hyper-V, Reports → ElementExecutive rollup
AdminAdmin (nodes, links, Dude import), TeamsQoE
ExecOverview health, Reports → Executive, CapacityTopology drill

SSO is Entra ID (demo-connected). Roles: Admin, NOC, Engineer, Exec, Read-only. This preview is not locked behind login. On the live server, SSO gates the same screens.

3. The language of the board

Learn this before you click anything. Same colors on map, alerts, reports, Hyper-V.

You seeMeansDo this
Green / UpResponding. Health typically ≥ 90.Nothing.
Yellow / DegradedUp, but not well (CPU, loss, pressure, crowded AP, cluster n-1).Investigate before it is an outage.
Red / Down / CriticalNot responding, circuit down, or Critical alert. Health 0.Incident.
Grey / unknownNot polled, or VM Saved/Off.Confirm it should be monitored.
Critical is red. Degraded is yellow. Red wakes someone. Yellow becomes red if you ignore it. Suppressed (dimmed) means the parent is already down or a maintenance window is on — do not page for Austin IPsec when DIA is dead.

Header, every page: live health, up / degraded / down. If health is not in the 90s, start at Overview or Alerting.

4. Health score

A node is not “up” just because ping works. Five slices, then a rollup:

  1. Availability — ICMP/SNMP. Zero if down, lost heartbeat, or VM Off/Saved.
  2. Performance — CPU, memory, Hyper-V CPU-wait, balloon.
  3. Hardware — temp, fans, PSUs; VHDX latency and host free RAM on Hyper-V.
  4. Response — ICMP vs that node’s own baseline (1→8 ms is worse than a branch that always sits at 30 ms).
  5. Interfaces — oper-down, high util, errors.

90–100 healthy · 70–89 degraded · below 70 or 0 treat as down. A SQL VM can be green on ICMP and yellow on Apps. Check the layer you care about: network vs hypervisor vs application.

5. The demo fabric

Unless someone imported a Dude file, you are on Northwind:

  • HQ Chicago — Palo Alto edge, Cisco access, Arista leaf, two 9800 WLCs, NAS, 2-node Hyper-V (HQ-HV-01 / 02) with CSV, SQL, DC, RDS, PBX.
  • Austin / Denver — Cisco ISR over IPsec. APs join HQ-WLC-9800-1 (FlexConnect).
  • DIA — two circuits: ISP-PE-CHI (DIA-1 2G) and ISP-PE-CHI-ZAYO (DIA-2 1G). Report DIA-1 first.
  • Wi-Fi world — NYC, São Paulo, London, Frankfurt, Dubai, Singapore, Tokyo, Sydney. Each has a site ISR and 2 ISPs (NYC / London / Singapore / Denver have a 3rd). APs CAPWAP to the two controllers over that local DIA.
  • M365 as a QoE target, not an SNMP box.

Inject from Overview or Hyper-V: WAN loss, AP down, host contention, cluster failover. They expire. That is the lab.

6. First ten minutes

Same path as the 15-minute card, written as a checklist. If you already did that, skip this.

  1. Overview — Critical/Degraded strip, site health, DIA, capacity days-to-exhaust.
  2. Click a site tile — Map opens already drilled.
  3. Map: Internet → firewall → switch → host/VM. Hover for inventory. Double-click dashboard.
  4. Open HQ-FW from Devices. Setup pills, probes, interfaces.
  5. NetPath — preset DIA (HQ-FW → ISP-PE). First red hop is where the user feels it.
  6. PerfStack — pick DIA, Core, SQL, or HV cluster. Do not build the first overlay.
  7. Hyper-V — two hosts, CSV owner, n-1 banner if a node is down.
  8. Reports → Element. It opens on DIA. Saved: DIA month, Site month, Worst 10.
  9. Admin — look, don’t rewrite production IPs. Reset demo restores Northwind.
  10. Alerting — Policies (six defaults), ack a minor, read the outbox.
Three clicks if you remember nothing else: Overview → Map drill → Reports Element (DIA).

7. Monitor

Overview

Operations home. Critical/Degraded strip first, then health, DIA, capacity, sites. Shortcuts sit under More. Run discovery is simulated. Inventory work lives in Admin.

Map

Tree from Internet down. Hover a node or a wire. Time travel plays snapshots back to the break.

NOC board

Six tiles: Sites, Open incidents, DIA, Cluster, Worst node, Alert age. Troubleshoot from Map / NetPath / the node, not from the TV.

Devices

Inventory columns: Status, Name, IP, Kind, Site, Health, Uptime. Node page has setup pills (ICMP/SNMP/WMI).

NetPath

Default path is DIA. Presets for Austin and Denver. Healthy firewall ping plus a red hop upstream is underlay/ISP, not LAN.

PerfStack

Presets: DIA, Core, SQL, HV cluster. Two lines spiking together are related until proven otherwise.

Wireless

APs, clients, noise. HQ-AP-FL1 at 27 stations vs baseline ~16 is a real seed example. Warehouse scanners stay on their VLAN.

Hyper-V

Two hosts, one cluster, CSV owner, witness, HA vs parked guests. Host: logical CPU, free RAM, VHDX, balloon. Guest: vCPU wait, assigned vs demand, heartbeat, state. Parked VMs must show “—”. When HV-01 is down, HV-02 yellow is correct (cluster n-1), not a second incident. SQL/DC/RDS also appear under Apps.

Experience

Off the sidebar. Application share and RTT still exist at /qoe if you need them. Pair with Traffic when “the internet is slow.”

8. Analyze

Traffic

NetFlow top talkers: who filled DIA, who is pounding CSV. Unknown UDP at the firewall is worth a look.

Configs

Running vs startup diff, CIS-style compliance. Austin shop NAT drift is the teaching example. Come here after a change window before you reload the box.

Apps

SQL (batch, page life, locks, log flush), AD (LDAP bind, replication), RDS (sessions, input delay). Page life falling while the VM looks fine is still a DBA incident — and you can prove the host was not starving it.

IPAM & UDT

Subnets, used/reserved/conflict IPs, and switch-port → MAC → user. “Who is on Gi1/0/8?” lives here. Demo includes a printer/laptop conflict on purpose.

Insight

PAN sessions, DP CPU, IKE tunnels with RTT, drop reasons, BGP. ICMP to the firewall is not “IKE is up.”

Capacity

Exhaustion dates. SQL RAM ~18 days should scare you; DIA ~47 days is the ISP conversation. Baselines: high-but-normal for that box is not an incident.

Events

Stream (syslog/traps) · Maintenance (toggle windows) · Dependencies (parent→child suppress) · Pollers + custom OIDs.

9. Operate

Alerting

Critical red, degraded yellow. Channels write to the outbox in demo — they do not leave the network. Ack clears the badge; history stays.

Reports

Element is the default. Filter DIA / routers / switches / … Read the headline paragraph first — that is why you ran it. Then current KPIs, history, this object’s outages only. Start DIA. Click a worst-node on Downtime to jump into its Element report.

Other tabs: Downtime (SLA, MTTR, MTBF), Patterns (heatmap, Tuesday 02:00, flapping, blast radius), Availability vs 99.95%, Executive (Monday email button), Capacity/Traffic.

Teams

Directory, role matrix, SSO connected/paused. Preview stays open.

Admin

Nodes (form editor — table is read-only so polling doesn’t steal focus), Links, Import. Merge is default (keeps Hyper-V). Replace wipes the fabric — training should not. Reset demo restores Northwind.

10. Playbook: working an incident

  1. Header — health not in the 90s? Down vs degraded.
  2. Alerting — newest Critical. Read it. Do not ack yet.
  3. Map, fabric — one site or the WAN? DIA red → don’t chase Austin.
  4. Events → Dependencies — suppressed? Work the parent.
  5. NetPath — first red hop is the fault domain.
  6. PerfStack — backup window vs ISP vs host pressure.
  7. VM? Hyper-V host wait / CSV / heartbeat, then Apps.
  8. WAN? Insight (IKE, BGP) + Traffic + Reports Element on DIA.
  9. Configs — commit in the last hour?
  10. Ack. Write the cause in manager language: “DIA loss 1.8%. Austin is downstream, not a branch fault.”
Inject WAN loss, then HV contention, then cluster failover, and walk this list each time. That is the lab.

11. Playbook: running a report

  1. Reports → Element.
  2. Chip DIA (or the ISR they named).
  3. Read the headline. If it answers the question, paste it into the ticket and stop.
  4. 30-day availability and this object’s outages — not the fabric log.
  5. “Every Tuesday?” → Patterns heatmap / recurring windows.
  6. SLA language → Downtime or Availability vs 99.95%.
  7. Email downtime report is the exec brief (demo outbox).

More data is there if you scroll. Do not send a 40-row table when the paragraph already said “two ISP events, 41 minutes, SLA breach.”

12. Importing from The Dude

Lattice is not Dude. Dude is an inventory source.

  1. Admin → Import.
  2. Leave mode on Merge unless you were told to replace.
  3. Drop the backup (XML / export / sqlite) or Load Dude sample.
  4. Read the toast: counts and warnings.
  5. Spot-check Map, a known IP, and that the Hyper-V cluster is still two hosts if you merged onto Northwind.

Ugly names (`device-12`) get renamed in Admin → Nodes. Lattice keeps the IP.

13. Calibrate (color → cause)

SymptomColorWhereTypical cause
DIA util 50s, loss 0.2%Green/yellowReports DIABackup window, not down
DIA loss > 1%Red alertAlerting + InsightISP / WAN inject
HV-01 wait, guests balloonYellowHyper-VHost contention inject
HV-01 down, SQL migratingHV-01 red, HV-02 yellowHyper-VFailover — one incident
AP 27+ clientsYellowWirelessOvercrowd inject
Austin down, DIA downAustin suppressedEvents depsWork DIA
SQL page life low, host fineApps yellowAppsDatabase, not compute
Saved VMGrey, metrics —Hyper-VParked — do not “fix RAM”

14. Demo vs live

CapabilityHereLive poller
Screens, health, maps, reportsProductSame
SNMP / ICMP / WinRMSimulated tickReal poll
Traps / syslog / NetFlowSimulatedReal receivers
Email / SMSOutbox previewSMTP / SMS gateway
SSODirectory, no login wallEntra SAML gate
Dude importParser + mergeSame parser
Inject buttonsTrainingDisable or keep staged

If a number looks too perfect, it is the simulator. You are learning the workflow.

15. Glossary

TermMeaning
DIADedicated Internet Access — HQ’s circuit to the ISP.
Health0–100 composite, not ping.
NetPathHop-by-hop path and quality.
PerfStackMulti-metric overlay on one time axis.
CSVCluster Shared Volume (Hyper-V).
CPU waitGuest ready but waiting on host CPU.
BalloonHost reclaiming guest RAM.
UDTPort / MAC / user.
UnDPExtra OID poller.
MTTR / MTBFMean time to restore / between failures.
SLA99.95% unless leadership says otherwise.
Merge vs ReplaceImport additively vs wipe the fabric.
SuppressedMuted because parent is down or maintenance is on.
You are trained enough to sit a shift when you can: (1) explain yellow vs red without a menu, (2) drill fabric → site → node, (3) run the incident playbook on injected WAN loss, (4) open Reports → Element → DIA and read the headline out loud, (5) merge a Dude sample without wiping Hyper-V, (6) tell a manager whether SQL is a VM problem or a database problem.