# Best practice for measuring and monitoring chef-client runs?

**URL:** <https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731>\
**Category:** Chef Infra (archive)\
**Created:** [September 17, 2014, 12:33am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731 "2014-09-17T00:33:06Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Augie\_Schwer](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/augie_schwer/32/246_2.png) [@Augie\_Schwer](https://discourse.chef.io/u/Augie_Schwer)\
**Post date:** [September 17, 2014, 12:33am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/1 "2014-09-17T00:33:06Z")

</div>

What are people using to monitor and measure their chef-client runs?

I would like to monitor for when chef-client runs fail on a node.

It would be nice to measure chef-client run times.

Is it safe to assume people are using handlers for both of these? What are  
some popular ways to accomplish these goals? Thanks!

–  
Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![jeffbyrnes](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/jeffbyrnes/32/1478_2.png) [@jeffbyrnes](https://discourse.chef.io/u/jeffbyrnes)\
**Post date:** [September 17, 2014, 12:38am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/2 "2014-09-17T00:38:49Z")

</div>

We use an email handler to report runs; primarily filtered for failed runs. Crude, but it works.

On Tue, Sep 16, 2014 at 8:33 PM, Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
wrote:

> ## What are people using to monitor and measure their chef-client runs? I would like to monitor for when chef-client runs fail on a node. It would be nice to measure chef-client run times. Is it safe to assume people are using handlers for both of these? What are some popular ways to accomplish these goals? Thanks!
> 
> Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![Mike](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/mike/32/18_2.png) [@Mike](https://discourse.chef.io/u/Mike)\
**Post date:** [September 17, 2014, 2:10am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/3 "2014-09-17T02:10:10Z")

</div>

> **[Chef Handler Datadog - Chef Supermarket](https://supermarket.getchef.com/tools/chef-handler-datadog)**
>
> Chef Handler Datadog

On Sep 16, 2014 8:39 PM, "Jeff Byrnes" [jeff@evertrue.com](mailto:jeff@evertrue.com) wrote:

> We use an email handler to report runs; primarily filtered for failed  
> runs. Crude, but it works.
> 
> On Tue, Sep 16, 2014 at 8:33 PM, Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
> wrote:
> 
> > What are people using to monitor and measure their chef-client runs?
> > 
> > I would like to monitor for when chef-client runs fail on a node.
> > 
> > It would be nice to measure chef-client run times.
> > 
> > Is it safe to assume people are using handlers for both of these? What  
> > are some popular ways to accomplish these goals? Thanks!
> > 
> > --  
> > Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![adamhjk](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/adamhjk/32/190_2.png) [@adamhjk](https://discourse.chef.io/u/adamhjk)\
**Post date:** [September 17, 2014, 2:12am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/4 "2014-09-17T02:12:55Z")

</div>

In addition, you can use the management console and analytics to get this  
data. Free under 25 nodes.  
On Sep 16, 2014 5:33 PM, "Augie Schwer" [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com) wrote:

> What are people using to monitor and measure their chef-client runs?
> 
> I would like to monitor for when chef-client runs fail on a node.
> 
> It would be nice to measure chef-client run times.
> 
> Is it safe to assume people are using handlers for both of these? What are  
> some popular ways to accomplish these goals? Thanks!
> 
> --  
> Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![Morgan\_Blackthorne](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/morgan_blackthorne/32/100_2.png) [@Morgan\_Blackthorne](https://discourse.chef.io/u/Morgan_Blackthorne)\
**Post date:** [September 17, 2014, 2:48am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/5 "2014-09-17T02:48:05Z")

</div>

We use Airbrake Handler to send the errors to hoptoad (which aggregates and  
emails). Haven't had time to dig into the run time analysis, not sure it  
matters to us at this point... complicated for us runs usually finish in  
about 1-2m, so that's plenty fine by me.

--  
~_~ StormeRider ~_~

"Every world needs its heroes [...] They inspire us to be better than we  
are. And they protect from the darkness that's just around the corner."

(from Smallville Season 6x1: "Zod")

On why I hate the phrase "that's so lame"... [http://bit.ly/Ps3uSS](http://bit.ly/Ps3uSS)

On Tue, Sep 16, 2014 at 5:33 PM, Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
wrote:

> What are people using to monitor and measure their chef-client runs?
> 
> I would like to monitor for when chef-client runs fail on a node.
> 
> It would be nice to measure chef-client run times.
> 
> Is it safe to assume people are using handlers for both of these? What are  
> some popular ways to accomplish these goals? Thanks!
> 
> --  
> Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![ppalanisamy](https://avatars.discourse-cdn.com/v4/letter/p/278dde/32.png) [@ppalanisamy](https://discourse.chef.io/u/ppalanisamy)\
**Post date:** [September 17, 2014, 2:50am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/6 "2014-09-17T02:50:35Z")

</div>

We use both email handler & Datadog handler. We were hit by a situation where there was a memory leak (with chef-client in daemon mode) which caused the handler also to fail without enough memory. We ended up fixing the memory leak and changed chef-client execution to task instead of service.

Thanks,  
Prakash

From: Mike [[mailto:miketheman@gmail.com](mailto:miketheman@gmail.com)]  
Sent: Wednesday, September 17, 2014 4:10 AM  
To: [chef@lists.opscode.com](mailto:chef@lists.opscode.com)  
Subject: [chef] Re: Re: Best practice for measuring and monitoring chef-client runs?

[https://supermarket.getchef.com/tools/chef-handler-datadog](https://supermarket.getchef.com/tools/chef-handler-datadog)  
On Sep 16, 2014 8:39 PM, “Jeff Byrnes” \<[jeff@evertrue.com](mailto:jeff@evertrue.com)[mailto:jeff@evertrue.com](mailto:jeff@evertrue.com)\> wrote:  
We use an email handler to report runs; primarily filtered for failed runs. Crude, but it works.

On Tue, Sep 16, 2014 at 8:33 PM, Augie Schwer \<[augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)[mailto:augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)\> wrote:  
What are people using to monitor and measure their chef-client runs?

I would like to monitor for when chef-client runs fail on a node.

It would be nice to measure chef-client run times.

Is it safe to assume people are using handlers for both of these? What are some popular ways to accomplish these goals? Thanks!

–  
Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us)[mailto:Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

Click here[https://www.mailcontrol.com/sr/MZbqvYs5QwJvpeaetUwhCQ==](https://www.mailcontrol.com/sr/MZbqvYs5QwJvpeaetUwhCQ==) to report this email as spam.

[![](http://www.sdl.com/Content/images/SDLlogo2014.png)  
www.sdl.com](http://www.sdl.com/?utm_source=Email&utm_medium=Email%2BSignature&utm_campaign=SDL%2BStandard%2BEmail%2BSignature)  
  

**SDL PLC confidential, all rights reserved.**

If you are not the intended recipient of this mail SDL requests and requires that you delete it without acting upon or copying any of its contents,  
and we further request that you advise us.  
  
SDL PLC is a public limited company registered in England and Wales.  
Registered number: 02675207.  
  
Registered address: Globe House, Clivemont Road, Maidenhead, Berkshire SL6 7DY, UK.

This message has been scanned for malware by Websense. [www.websense.com](http://www.websense.com)

---

<div class="post-metadata">

**Author:** ![Steffen\_Gebert1](https://avatars.discourse-cdn.com/v4/letter/s/e56c9b/32.png) [@Steffen\_Gebert1](https://discourse.chef.io/u/Steffen_Gebert1)\
**Post date:** [September 17, 2014, 6:03am UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/7 "2014-09-17T06:03:58Z")

</div>

Jumping into the "we do.." postings: We send chef-client statistics to  
zabbix using a report handler:

- success
- elapsed\_time
- start\_time
- end\_time
- all\_resources\_num
- updated\_resource num

I think that should be pretty easy to adapt to whatever monitoring  
system you use.

Yours  
Steffen

Links:

> <https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb>

[https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb)

On 17/09/14 02:33, Augie Schwer wrote:

> What are people using to monitor and measure their chef-client runs?
> 
> I would like to monitor for when chef-client runs fail on a node.
> 
> It would be nice to measure chef-client run times.
> 
> Is it safe to assume people are using handlers for both of these? What are  
> some popular ways to accomplish these goals? Thanks!

---

<div class="post-metadata">

**Author:** ![Mark\_Mzyk\_OLD](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/mark_mzyk_old/32/197_2.png) [@Mark\_Mzyk\_OLD](https://discourse.chef.io/u/Mark_Mzyk_OLD)\
**Post date:** [September 17, 2014, 4:37pm UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/8 "2014-09-17T16:37:10Z")

</div>

The report handler that supplies the data from the client run to the  
Chef server reporting add on is open source, so it could be used and/or  
built off of, if you didn't want to use the pre-built Chef add ons.

It's here in the client:

> <https://github.com/chef/chef/blob/main/lib/chef/resource_reporter.rb>

- Mark Mzyk

> Steffen Gebert [mailto:st+gmane@st-g.de](mailto:st+gmane@st-g.de)  
> September 17, 2014 at 2:03 AM  
> Jumping into the "we do.." postings: We send chef-client statistics to  
> zabbix using a report handler:
> 
> - success
> - elapsed\_time
> - start\_time
> - end\_time
> - all\_resources\_num
> - updated\_resource num
> 
> I think that should be pretty easy to adapt to whatever monitoring  
> system you use.
> 
> Yours  
> Steffen
> 
> Links:  
> [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb)  
> [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb)
> 
> Augie Schwer [mailto:augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
> September 16, 2014 at 8:33 PM  
> What are people using to monitor and measure their chef-client runs?
> 
> I would like to monitor for when chef-client runs fail on a node.
> 
> It would be nice to measure chef-client run times.
> 
> Is it safe to assume people are using handlers for both of these? What  
> are some popular ways to accomplish these goals? Thanks!
> 
> --  
> Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

---

<div class="post-metadata">

**Author:** ![DV1](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/dv1/32/186_2.png) [@DV1](https://discourse.chef.io/u/DV1)\
**Post date:** [September 17, 2014, 6:35pm UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/9 "2014-09-17T18:35:02Z")

</div>

We have a custom Rails app that acts as handler for chef-client. Here's  
what dashboard looks like: [http://i.imgur.com/sR4UCWC.png](http://i.imgur.com/sR4UCWC.png)

We also have an automated task that runs "knife status" and reports on any  
hosts that haven't checked in for a while.

On Wed, Sep 17, 2014 at 9:37 AM, Mark Mzyk [mmzyk@getchef.com](mailto:mmzyk@getchef.com) wrote:

> The report handler that supplies the data from the client run to the Chef  
> server reporting add on is open source, so it could be used and/or built  
> off of, if you didn't want to use the pre-built Chef add ons.
> 
> It's here in the client:  
> [https://github.com/opscode/chef/blob/master/lib/chef/resource\_reporter.rb](https://github.com/opscode/chef/blob/master/lib/chef/resource_reporter.rb)
> 
> - Mark Mzyk
> 
> - success
> 
> - elapsed\_time
> 
> - start\_time
> 
> - end\_time
> 
> - all\_resources\_num
> 
> - updated\_resource num
> 
> I think that should be pretty easy to adapt to whatever monitoring  
> system you use.
> 
> Yours  
> Steffen
> 
> Links:
> 
> [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb)
> 
> [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb)
> 
> Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
> September 16, 2014 at 8:33 PM  
> What are people using to monitor and measure their chef-client runs?
> 
> I would like to monitor for when chef-client runs fail on a node.
> 
> It would be nice to measure chef-client run times.
> 
> Is it safe to assume people are using handlers for both of these? What are  
> some popular ways to accomplish these goals? Thanks!
> 
> --  
> Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)

--  
Best regards, Dmitriy V.

---

<div class="post-metadata">

**Author:** ![Tiago\_Cruz](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/tiago_cruz/32/49_2.png) [@Tiago\_Cruz](https://discourse.chef.io/u/Tiago_Cruz)\
**Post date:** [September 17, 2014, 7:04pm UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/10 "2014-09-17T19:04:45Z")

</div>

I have a simple check on nagios like this

...  
HOURS=2  
SECONDS=$(expr $HOURS \* 60 \* 60)  
OHAI\_TIME="$(expr $(date +%s) - $SECONDS)"

SEARCH="$(knife search node "ohai\_time:[\* TO $OHAI\_TIME] AND  
chef\_environment:production")"  
...

On Wed, Sep 17, 2014 at 3:35 PM, DV [vindimy@gmail.com](mailto:vindimy@gmail.com) wrote:

> We have a custom Rails app that acts as handler for chef-client. Here's  
> what dashboard looks like: [http://i.imgur.com/sR4UCWC.png](http://i.imgur.com/sR4UCWC.png)
> 
> We also have an automated task that runs "knife status" and reports on any  
> hosts that haven't checked in for a while.
> 
> On Wed, Sep 17, 2014 at 9:37 AM, Mark Mzyk [mmzyk@getchef.com](mailto:mmzyk@getchef.com) wrote:
> 
> > The report handler that supplies the data from the client run to the Chef  
> > server reporting add on is open source, so it could be used and/or built  
> > off of, if you didn't want to use the pre-built Chef add ons.
> > 
> > It's here in the client:  
> > [https://github.com/opscode/chef/blob/master/lib/chef/resource\_reporter.rb](https://github.com/opscode/chef/blob/master/lib/chef/resource_reporter.rb)
> > 
> > - Mark Mzyk
> > 
> > - success
> > 
> > - elapsed\_time
> > 
> > - start\_time
> > 
> > - end\_time
> > 
> > - all\_resources\_num
> > 
> > - updated\_resource num
> > 
> > I think that should be pretty easy to adapt to whatever monitoring  
> > system you use.
> > 
> > Yours  
> > Steffen
> > 
> > Links:
> > 
> > [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb)
> > 
> > [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb)
> > 
> > Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
> > September 16, 2014 at 8:33 PM  
> > What are people using to monitor and measure their chef-client runs?
> > 
> > I would like to monitor for when chef-client runs fail on a node.
> > 
> > It would be nice to measure chef-client run times.
> > 
> > Is it safe to assume people are using handlers for both of these? What  
> > are some popular ways to accomplish these goals? Thanks!
> > 
> > --  
> > Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)
> 
> --  
> Best regards, Dmitriy V.

--  
-- Tiago Cruz

---

<div class="post-metadata">

**Author:** ![Augie\_Schwer](https://sea2.discourse-cdn.com/flex016/user_avatar/discourse.chef.io/augie_schwer/32/246_2.png) [@Augie\_Schwer](https://discourse.chef.io/u/Augie_Schwer)\
**Post date:** [September 17, 2014, 11:52pm UTC](https://discourse.chef.io/t/best-practice-for-measuring-and-monitoring-chef-client-runs/5731/11 "2014-09-17T23:52:55Z")

</div>

Thanks everyone, that is all very helpful.

On Wed, Sep 17, 2014 at 12:04 PM, Tiago Cruz [tiago.tuxkiller@gmail.com](mailto:tiago.tuxkiller@gmail.com)  
wrote:

> I have a simple check on nagios like this
> 
> ...  
> HOURS=2  
> SECONDS=$(expr $HOURS \* 60 \* 60)  
> OHAI\_TIME="$(expr $(date +%s) - $SECONDS)"
> 
> SEARCH="$(knife search node "ohai\_time:[\* TO $OHAI\_TIME] AND  
> chef\_environment:production")"  
> ...
> 
> On Wed, Sep 17, 2014 at 3:35 PM, DV [vindimy@gmail.com](mailto:vindimy@gmail.com) wrote:
> 
> > We have a custom Rails app that acts as handler for chef-client. Here's  
> > what dashboard looks like: [http://i.imgur.com/sR4UCWC.png](http://i.imgur.com/sR4UCWC.png)
> > 
> > We also have an automated task that runs "knife status" and reports on  
> > any hosts that haven't checked in for a while.
> > 
> > On Wed, Sep 17, 2014 at 9:37 AM, Mark Mzyk [mmzyk@getchef.com](mailto:mmzyk@getchef.com) wrote:
> > 
> > > The report handler that supplies the data from the client run to the  
> > > Chef server reporting add on is open source, so it could be used and/or  
> > > built off of, if you didn't want to use the pre-built Chef add ons.
> > > 
> > > It's here in the client:  
> > > [https://github.com/opscode/chef/blob/master/lib/chef/resource\_reporter.rb](https://github.com/opscode/chef/blob/master/lib/chef/resource_reporter.rb)
> > > 
> > > - Mark Mzyk
> > > 
> > > - success
> > > 
> > > - elapsed\_time
> > > 
> > > - start\_time
> > > 
> > > - end\_time
> > > 
> > > - all\_resources\_num
> > > 
> > > - updated\_resource num
> > > 
> > > I think that should be pretty easy to adapt to whatever monitoring  
> > > system you use.
> > > 
> > > Yours  
> > > Steffen
> > > 
> > > Links:
> > > 
> > > [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/recipes/chef-client.rb)
> > > 
> > > [https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb](https://github.com/TYPO3-cookbooks/zabbix-custom-checks/blob/master/templates/default/chef-client/chef-client-handler.rb)
> > > 
> > > Augie Schwer [augie.schwer@gmail.com](mailto:augie.schwer@gmail.com)  
> > > September 16, 2014 at 8:33 PM  
> > > What are people using to monitor and measure their chef-client runs?
> > > 
> > > I would like to monitor for when chef-client runs fail on a node.
> > > 
> > > It would be nice to measure chef-client run times.
> > > 
> > > Is it safe to assume people are using handlers for both of these? What  
> > > are some popular ways to accomplish these goals? Thanks!
> > > 
> > > --  
> > > Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)
> > 
> > --  
> > Best regards, Dmitriy V.
> 
> --  
> -- Tiago Cruz

--  
Augie Schwer - [Augie@Schwer.us](mailto:Augie@Schwer.us) - [http://schwer.us](http://schwer.us)
