<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <generator uri="http://jekyllrb.com" version="3.10.0">Jekyll</generator>
  
  
  <link href="https://anthonylouisdagostino.com/feed.xml" rel="self" type="application/atom+xml" />
  <link href="https://anthonylouisdagostino.com/" rel="alternate" type="text/html" />
  <updated>2026-01-02T04:26:03+00:00</updated>
  <id>https://anthonylouisdagostino.com/</id>

  
    <title type="html">Anthony Louis D’Agostino, PhD</title>
  

  
    <subtitle>Research Economist, Mathematica</subtitle>
  

  
    <author>
        <name>Anthony Louis D&apos;Agostino</name>
      
      
    </author>
  

  
  
    <entry>
      
      <title type="html">Dose of Data - Jan 01, 2026</title>
      
      
      <link href="https://anthonylouisdagostino.com/dod-20260101/" rel="alternate" type="text/html" title="Dose of Data - Jan 01, 2026" />
      
      <published>2026-01-01T00:00:00+00:00</published>
      <updated>2026-01-01T00:00:00+00:00</updated>
      <id>https://anthonylouisdagostino.com/dose-of-data</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/dod-20260101/">&lt;ul&gt;
  &lt;li&gt;6.2% - &lt;a href=&quot;https://www.bls.gov/lau/&quot; target=&quot;_blank&quot;&gt;DC’s unemployment rate&lt;/a&gt; in September 2025 (seasonally adjusted)&lt;/li&gt;
  &lt;li&gt;69 - &lt;a href=&quot;https://homicides.news.baltimoresun.com/?range=2025&quot; target=&quot;_blank&quot;&gt;year on year reduction in homicides&lt;/a&gt; in Baltimore in 2025&lt;/li&gt;
  &lt;li&gt;$480M - the amount in health funding the &lt;a href=&quot;https://apnews.com/article/ivory-coast-us-health-deal-usaid-45ae5d2b876be76d7a0d62f3c523cec2&quot; target=&quot;_blank&quot;&gt;US has committed to the Ivory Coast&lt;/a&gt;, covering issue areas including HIV, malaria, maternal and child health, and health security, matched by an expected $292M from the Ivory Coast by 2030&lt;/li&gt;
  &lt;li&gt;1st - the birthday recently celebrated by &lt;a href=&quot;https://www.news18.com/india/tarmeem-turns-one-the-story-behind-indias-first-gene-edited-sheep-ws-kl-9797918.html&quot; target=&quot;_blank&quot;&gt;Tarmeem&lt;/a&gt;, India’s first CRISPR gene-edited sheep&lt;/li&gt;
&lt;/ul&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="data" />
      
        <category term="DC" />
      
        <category term="policy" />
      
        <category term="research" />
      

      

      
        <summary type="html">6.2% - DC’s unemployment rate in September 2025 (seasonally adjusted) 69 - year on year reduction in homicides in Baltimore in 2025 $480M - the amount in health funding the US has committed to the Ivory Coast, covering issue areas including HIV, malaria, maternal and child health, and health security, matched by an expected $292M from the Ivory Coast by 2030 1st - the birthday recently celebrated by Tarmeem, India’s first CRISPR gene-edited sheep</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Beauty in Drainage Coefficient Values</title>
      
      
      <link href="https://anthonylouisdagostino.com/ddc-12-2025/" rel="alternate" type="text/html" title="Beauty in Drainage Coefficient Values" />
      
      <published>2025-12-26T00:00:00+00:00</published>
      <updated>2025-12-26T00:00:00+00:00</updated>
      <id>https://anthonylouisdagostino.com/ddc-roc</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/ddc-12-2025/">&lt;p&gt;&lt;img src=&quot;/imgs/RepCongo_and_DRC_DDC_Values_from_GESD_Raster.png&quot; alt=&quot;image&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Tributaries along the Congo River, as seen from the perspective of deep drainage coefficient (DDC) values modeled in &lt;a href=&quot;https://agupubs.onlinelibrary.wiley.com/doi/full/10.1002/2013MS000293&quot;&gt;Shangguan et al. (2014)&lt;/a&gt; at 30 arc-second resolution globally.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="art" />
      
        <category term="soil" />
      
        <category term="ddc" />
      
        <category term="Africa" />
      

      

      
        <summary type="html"></summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">The Primacy of Payback Periods for Energy Efficiency</title>
      
      
      <link href="https://anthonylouisdagostino.com/energypayback-02-2024/" rel="alternate" type="text/html" title="The Primacy of Payback Periods for Energy Efficiency" />
      
      <published>2024-02-22T21:12:23+00:00</published>
      <updated>2024-02-22T21:12:23+00:00</updated>
      <id>https://anthonylouisdagostino.com/incompetent-disingenuous</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/energypayback-02-2024/">&lt;p&gt;A &lt;a href=&quot;https://www.washingtonpost.com/climate-environment/2024/02/21/homebuilders-energy-efficiency-climate/&quot;&gt;WaPo article&lt;/a&gt; from yesterday details state-level efforts often spearheaded by developers, to roll back or slow down building code reforms that would increase energy efficiency improvements in new housing. A key dividing line between developers and EE advocates is the upfront premium for EE improvements like better insulation, triple glazed windows, and electrical systems that would facilitate EV charger upgrades. While developers quoted in the article claim $20K (perhaps a 10% premium on home value), EE advocates argue it’s closer to $6K. &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;One developer who is building EE cottages in North Carolina shared this thought-provoking statement: &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;“We do need to wrestle with the issue of cost, but it strikes me funny that we’re measuring improvements to houses by this simple payback calculation,” he said. “Nobody is asking you what the payback is on your fancy cabinets or flooring. But energy efficiency always comes down to that debate.” &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the immediate answer is that cabinets and flooring are of course still subject to payback calculations, but the math pencils out for them because it makes the owner happy and their &lt;em&gt;willingness to pay&lt;/em&gt; is what seals the purchase. EE investments of course can also lead to occupant happiness (or misery, especially for their absence), but there’s a solid counterfactual with knowable and recurring costs against which lump-sum, upfront expenses can be compared. Just like there’s some cash-in-hand value at which a homebuyer would be indifferent between a home with premium &lt;a href=&quot;https://puustelliusa.com/&quot;&gt;Puustelli&lt;/a&gt; cabinets and a more humble alternative replete with Hampton Bay cabinetry with $X thousand savings, homeowners at some breakpoint discount rate should be equally satisfied between a more expensive but efficient home and a cheaper yet draftier version. &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;So where do you stand? Assume you’re shopping for a home in a new state and don’t have any prior utility bills to evaluate monthly energy costs. A first point of contact might be the DOE’s &lt;a href=&quot;https://www.eia.gov/consumption/residential/data/2020/index.php?view=state&quot;&gt;Residential Energy Consumption Survey&lt;/a&gt; which is a quasi-quinquennial, national survey of nearly 20,000 households. The latest data (2020) for NC indicates an average annual energy bill of $1,731, totaling up expenses from space heating, air conditioning, water heating, and refrigeration. Assume that EE investments could shave off 20%, or $340 in annual savings. Ten years of savings with a 0% discount rate (for simplicity) are only slightly north of $3,000, and that number is decently higher than the true present value discounted at 2-3%. Holding to the side that EE improvements and energy code compliance costs do vary by state, we can alternatively look at some of the higher energy cost states and reassess when and how EE premia pay for themselves in the long run. CT is a good contender here, with an annual bill of $2,800. &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Now one thing to keep in mind when reviewing the RECS data is these are all pre-pandemic responses and therefore exclude the additional energy costs associated from working at home. For households with any members in WFH status, space heating and air conditioning totals are likely to be substantially higher than the 2020 values where houses are vacant during the day and thermostats set accordingly. With many working professionals still WFH for at least part of the week, the payback period for EE improvements would shorten even further.   &lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="building code" />
      
        <category term="energy efficiency" />
      
        <category term="housing" />
      
        <category term="affordability" />
      

      

      
        <summary type="html">A WaPo article from yesterday details state-level efforts often spearheaded by developers, to roll back or slow down building code reforms that would increase energy efficiency improvements in new housing. A key dividing line between developers and EE advocates is the upfront premium for EE improvements like better insulation, triple glazed windows, and electrical systems that would facilitate EV charger upgrades. While developers quoted in the article claim $20K (perhaps a 10% premium on home value), EE advocates argue it’s closer to $6K.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Book Review: John Doerr and Ryan Panchadsaram. *Speed &amp;amp; Scale: A Global Action Plan for Solving Our Climate Crisis Now* (Penguin Business 2021)</title>
      
      
      <link href="https://anthonylouisdagostino.com/speed&scale-01-2023/" rel="alternate" type="text/html" title="Book Review: John Doerr and Ryan Panchadsaram. *Speed &amp; Scale: A Global Action Plan for Solving Our Climate Crisis Now* (Penguin Business 2021)" />
      
      <published>2023-01-30T07:12:23+00:00</published>
      <updated>2023-01-30T07:12:23+00:00</updated>
      <id>https://anthonylouisdagostino.com/speed-&amp;-scale</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/speed&amp;scale-01-2023/">&lt;p&gt;If you asked for a single book to provide a comprehensive blueprint for how we might achieve net zero by 2050 and minimize the chances of facing devastating climate change, this might have been that book. Doerr and Panchadsaram, both of Kleiner Perkins, start &lt;i&gt;Speed &amp;amp; Scale&lt;/i&gt; with a sector by sector play to cancel out ~60 Gt/year of CO&lt;sub&gt;2&lt;/sub&gt;-eq emissions. First, electrify transportation, decarbonize the grid, “fix food”, protect nature, and remove carbon from the atmosphere (and the oceans), this last part attached to a lofty 10 Gt target. The targets and timetables adopt an Objectives and Key Results (OKR) framework which the authors emphasize in the introduction as critical for success. Standard “you can only manage what you measure” messaging.&lt;/p&gt;

&lt;p&gt;While the book is meant to be a banquet table of solutions wrapped in an aura of innovation and smothered in magisterial scope, the &lt;i&gt;plan&lt;/i&gt; is primarily a technocratic plug-in-play alteration of our current existence. Cars? Replace with EVs. Electricity? More renewables, more battery storage. Supply chains? Better monitoring and land set asides.&lt;/p&gt;

&lt;p&gt;Towards the end of the book, Laurene Powell Jobs provides one of the few worthwhile quotations: “You have to address everything at the same time.”&lt;/p&gt;

&lt;p&gt;I had hoped Doerr and Panchadsaram might reflect on how the climate crisis could be an opportunity to reconfigure society in a way that simultaneously forestalls climate change while also fixing our other ills: gun violence, obesity, suicidality,  extreme partisanship…&lt;/p&gt;

&lt;p&gt;That holism peaks out at times, like in advocating for universal education and shrinking gender gaps in India through programs like Educate Girls.&lt;/p&gt;

&lt;p&gt;But urbanism and a restructuring of values and social relations was simply not on this banquet table. A lot of ink was spent on EVs and battery technology and Rivian and Teslas — not even a page was dedicated to bikes, e-bikes, and improved walkability. Classic technocratic plug-and-play. Current ICE cars bad, just replace with EVs. Meat? Ethan Brown of Beyond Meat fame has a call-out box and there’s some discussion about the carbon intensity of meat consumption, but there’s certainly no investment plan to convince the masses to forego the flesh. Examples like these strike me as big failings of the book. For the authors, it feels like the problem is one of hardware version, not values. If we can just upgrade to the next version, then humanity will be saved…&lt;/p&gt;

&lt;p&gt;This book is frustrating because it’s described through the prism of Kleiner Perkins investments. Sure, it’s definitely helpful to roll back the curtain on failed investments or the signals that investors use to assess the viability of new technologies, like renewables in the 2000s and “clean-tech” currently. But a lot of the changes that need to happen are tangential to investments – it’s about pressure campaigns on governments local and federal, and shifting public consciousness about what constitutes a full life.&lt;/p&gt;

&lt;p&gt;In short, the vision in &lt;i&gt;Speed &amp;amp; Scale&lt;/i&gt; is that if we can innovate, invest, and shave down the green premium – the added that the ‘sustainable’ option has over the existing, unsustainable alternative – then we can solve our way out of this problem. If $1.7 trillion can be invested per year over the next 20 years, then we can have matured the necessary technologies.&lt;/p&gt;

&lt;p&gt;And yet, in spite of all the advances in the past decades, 2022 still saw the largest annual amount of GHG emissions ever.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="decarbonization" />
      
        <category term="book review" />
      
        <category term="Doerr" />
      
        <category term="climate change" />
      
        <category term="sustainability" />
      

      

      
        <summary type="html">If you asked for a single book to provide a comprehensive blueprint for how we might achieve net zero by 2050 and minimize the chances of facing devastating climate change, this might have been that book. Doerr and Panchadsaram, both of Kleiner Perkins, start Speed &amp;amp; Scale with a sector by sector play to cancel out ~60 Gt/year of CO2-eq emissions. First, electrify transportation, decarbonize the grid, “fix food”, protect nature, and remove carbon from the atmosphere (and the oceans), this last part attached to a lofty 10 Gt target. The targets and timetables adopt an Objectives and Key Results (OKR) framework which the authors emphasize in the introduction as critical for success. Standard “you can only manage what you measure” messaging.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Bounding Boxes for all US Counties</title>
      
      
      <link href="https://anthonylouisdagostino.com/bounding-boxes-for-all-us-counties/" rel="alternate" type="text/html" title="Bounding Boxes for all US Counties" />
      
      <published>2021-11-05T07:12:23+00:00</published>
      <updated>2021-11-05T07:12:23+00:00</updated>
      <id>https://anthonylouisdagostino.com/bounding-boxes-for-all-us-counties</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/bounding-boxes-for-all-us-counties/">&lt;p&gt;&lt;img src=&quot;/wp-content/uploads/2021/11/county_boundingboxes.png&quot; alt=&quot;image&quot; /&gt;
A post from several years back contained the &lt;a href=&quot;/bounding-boxes-for-all-us-states/&quot;&gt;bounding box coordinates of all US states&lt;/a&gt; and has been one of the more viewed pages on this site. Unfortunately, if your area of interest is below the state-level, these bounding boxes may only get you part of the way to your destination. Why waste time expanding a geographic search to areas beyond your narrow AOI?&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;If you’re running county-level analyses and need the latitude and longitude bounding box endpoints, then this is the table for you. These values were generated using the TIGER/Line 2021 shapefiles based on Census 2020 geographies, which you can grab from the &lt;a href=&quot;https://www2.census.gov/geo/tiger/TIGER2021/&quot;&gt;Census FTP site&lt;/a&gt;.&lt;br /&gt;
&lt;br /&gt;
Need to construct bounding boxes for some geometry collection of your choice? This &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sf/tidyverse&lt;/code&gt; chunk might be a good start:&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;library(tidyverse)
library(sf)
us.counties &amp;lt;- st_read(&quot;tl_2021_us_county.shp&quot;, stringsAsFactors = FALSE)
bb.rbind.sf &amp;lt;- split(us.counties, 1:nrow(us.counties)) %&amp;gt;%
                map( ~ st_bbox(.x) %&amp;gt;%
                st_as_sfc() %&amp;gt;%
                as_tibble()) %&amp;gt;%
                do.call(&quot;rbind&quot;, . ) %&amp;gt;%  
                st_as_sf()
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;div class=&quot;is-layout-flow wp-block-group&quot;&gt;&lt;div class=&quot;wp-block-group__inner-container&quot;&gt;&lt;/div&gt;&lt;/div&gt;
&lt;script src=&quot;https://gist.github.com/a8dx/7e550680f7ea6a68f20da00e21d7ce9b.js&quot;&gt;&lt;/script&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="Data" />
      
        <category term="Geography" />
      
        <category term="GIS" />
      
        <category term="Visualization" />
      

      

      
        <summary type="html">A post from several years back contained the bounding box coordinates of all US states and has been one of the more viewed pages on this site. Unfortunately, if your area of interest is below the state-level, these bounding boxes may only get you part of the way to your destination. Why waste time expanding a geographic search to areas beyond your narrow AOI?</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Tip for Installing Orfeo Toolbox Plugin for QGIS on MacOS</title>
      
      
      <link href="https://anthonylouisdagostino.com/tip-for-installing-orfeo-toolbox-plugin-for-qgis-on-macos/" rel="alternate" type="text/html" title="Tip for Installing Orfeo Toolbox Plugin for QGIS on MacOS" />
      
      <published>2019-08-16T02:47:08+00:00</published>
      <updated>2019-08-16T02:47:08+00:00</updated>
      <id>https://anthonylouisdagostino.com/tip-for-installing-orfeo-toolbox-plugin-for-qgis-on-macos</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/tip-for-installing-orfeo-toolbox-plugin-for-qgis-on-macos/">&lt;p&gt;&lt;img src=&quot;/wp-content/uploads/2019/08/logo.png&quot; alt=&quot;image&quot; /&gt;&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Last weekend I spent longer than I care to admit trying to get &lt;a href=&quot;https://www.orfeo-toolbox.org/CookBook/index_TOC.html&quot; target=&quot;_blank&quot;&gt;Orfeo Toolbox&lt;/a&gt; (OTB) to play nicely with QGIS on MacOS High Sierra. I tried QGIS 3.4, the stable version, and I tried 3.8, the latest version, and scoured every forum and installation tutorial I could find for activating the tools and having them appear in QGIS’ Processing Toolbox. Once they appear, ensuring they actually work is a separate issue. One piece of advice seems to have made all the difference.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;So, I want to shout out Juan Ramón Selva for pulling me from the morass with what appears to be the most important detail that’s not adequately emphasized in any OTB installation instructions.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;When you extract &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;OTB-6.6.x-Darwin64.run&lt;/code&gt; that was downloaded from the OTB site in Terminal, it’ll create a folder entitled “OTB-6.6.x-Darwin64.” You may have read in some fora that OTB acts up if the target folder includes non-alphanumeric characters like hyphens or spaces, and have been tempted to manually rename the folder to something innocuous like “OTB.” As Juan points out &lt;a href=&quot;https://gitlab.orfeo-toolbox.org/orfeotoolbox/otb/issues/1745&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt;, that’s a no-go. Instead, you need to specify any differently-named target folder like “OTB”, as explained in his Oct 19, 2018 post.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;Once that was sorted out, I went back and revised the remaining installation steps:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Add the plugin and &lt;a href=&quot;https://gitlab.orfeo-toolbox.org/orfeotoolbox/qgis-otb-plugin/blob/master/README.md&quot; target=&quot;_blank&quot;&gt;specify the Name and URL&lt;/a&gt; in QGIS’ plugin manager.&lt;/li&gt;
  &lt;li&gt;Properly identify the OTB and OTB application folders, as described in the Open processing settings section &lt;a href=&quot;https://www.orfeo-toolbox.org/CookBook/QGIS-interface.html&quot; target=&quot;_blank&quot;&gt;here&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="computing" />
      
        <category term="GIS" />
      
        <category term="Orfeo" />
      

      

      
        <summary type="html"></summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Cleaning Berkeley Earth’s BEST Gridded Daily Temperature Data</title>
      
      
      <link href="https://anthonylouisdagostino.com/cleaning-berkeley-earths-best-gridded-daily-temperature-data/" rel="alternate" type="text/html" title="Cleaning Berkeley Earth&apos;s BEST Gridded Daily Temperature Data" />
      
      <published>2018-12-12T09:33:41+00:00</published>
      <updated>2018-12-12T09:33:41+00:00</updated>
      <id>https://anthonylouisdagostino.com/cleaning-berkeley-earths-best-gridded-daily-temperature-data</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/cleaning-berkeley-earths-best-gridded-daily-temperature-data/">&lt;p&gt;You may have recently seen air quality maps produced by the &lt;a href=&quot;http://berkeleyearth.org/&quot;&gt;Berkeley Earth&lt;/a&gt; group, especially in the wake of the &lt;a href=&quot;https://www.sfgate.com/california-wildfires/article/Camp-Fire-Death-toll-rises-to-86-after-13458956.php&quot;&gt;horrific Camp Fire&lt;/a&gt; whose death toll now exceeds 80. For example, &lt;a href=&quot;http://berkeleyearth.org/air-quality-real-time-map/&quot;&gt;here’s&lt;/a&gt; their real-time visualization of PM2.5 concentrations.&lt;/p&gt;

&lt;p&gt;For several years, the Berkeley Earth team had primarily worked on producing gridded weather datasets with solid historical coverage, going back to 1850 in some cases. I’ve been particularly interested in their work given the &lt;a href=&quot;http://berkeleyearth.org/data/&quot;&gt;global daily temperature datasets they have on tap&lt;/a&gt; – which are publicly available! This is all very exciting, given that many of the global datasets available, like &lt;a href=&quot;https://www.esrl.noaa.gov/psd/data/gridded/data.UDel_AirT_Precip.html&quot;&gt;Willmott and Matsuura (UDEL)&lt;/a&gt; and &lt;a href=&quot;https://crudata.uea.ac.uk/cru/data/hrg/&quot;&gt;CRU&lt;/a&gt;, are only at monthly resolution. However, there’s something to be desired in making their offerings truly accessible to researchers:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Their netCDF files are packaged with “temperature anomaly” and “climatology” layers, but no “temperature” layer. This necessitates a bit of wrangling on your part to overcome some inherent dimensional consistency issues to create a temperature object. It would’ve struck me to provide the temperature values, and allow researchers to determine which climatology period they want to construct, not the reverse.&lt;/li&gt;
  &lt;li&gt;For example, the climatology is a 365-day stack. Great, except that the anomalies data is based on true dates and includes leap years. There’s not an immediate way to sum the two into temperature estimates. If you ignore the mismatch and are performing analysis that spans many decades, by the end of your time-series you’ll be off by several weeks (amateurish!).&lt;/li&gt;
  &lt;li&gt;The .nc’s are stripped of an informative time axis (see below figure). Instead, the files have date_number, day, month, and year attributes. Why couldn’t they have provided an out of the box date object, and left it to researchers to parse the month and year when necessary?&lt;/li&gt;
  &lt;li&gt;As a result, you wouldn’t even know whether leap dates are included unless you manually inspected the layer count for a decadal file.&lt;/li&gt;
  &lt;li&gt;And lastly, the absence of that time axis means you’re not able to discern dates when visually inspecting the data in a viewer like &lt;a href=&quot;https://www.giss.nasa.gov/tools/panoply/download/&quot;&gt;Panoply&lt;/a&gt;. If you wanted to see what the temperature [anomalies] were for Jan 12, 2013, then you’ll have to manually count how many days passed since the Jan 1, 2010 start date of that decadal extract.&lt;/li&gt;
&lt;/ul&gt;

&lt;figure class=&quot;wp-block-image&quot;&gt;![](https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/12/best_data.png?resize=712%2C520&amp;amp;ssl=1)&lt;figcaption&gt;BEST temperature data does not come packaged with a useful time axis  
&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;I’ve spent a fair amount of time trying to address these issues, and wanted to share &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;R&lt;/code&gt; code so that you don’t have to recreate the process. The only packages you’ll need to run this are &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;raster&lt;/code&gt; and &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tidyverse&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;In brief, this function 1.) reads in a BEST .nc file, which you may or may not have previously spatial subsetted, 2.) creates a stacked climatology that properly accounts for leap years by duplicating Feb 28 climatology values for Feb 29, and 3.) outputs temperature estimates with a date explicit axis.&lt;/p&gt;

&lt;figure class=&quot;wp-block-embed is-type-rich&quot;&gt;&lt;div class=&quot;wp-block-embed__wrapper&quot;&gt;&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;&lt;div class=&quot;gist&quot; id=&quot;gist93531237&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-clean_best_data-r&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-r  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt; This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters. [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars) &lt;/span&gt;&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;&lt;/template&gt;&lt;template class=&quot;js-line-alert-template&quot;&gt; &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;&lt;/span&gt;&lt;/template&gt; |  |  |
|---|---|
|  | &lt;span class=&quot;pl-en&quot;&gt;returnBESTtemp&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;function&lt;/span&gt;(&lt;span class=&quot;pl-smi&quot;&gt;tempFile&lt;/span&gt;) { |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; tempFile: complete path to a netcdf BEST temperature file of daily records \[need not be geographic subset\]&lt;/span&gt; |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; brick(&lt;span class=&quot;pl-smi&quot;&gt;tempFile&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;var&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;climatology&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; brick(&lt;span class=&quot;pl-smi&quot;&gt;tempFile&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;var&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;temperature&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — deal with absence of leap year in anomaly series&lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;ave.dates&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; as\_date(as.numeric(unlist(lapply(strsplit(names(&lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;), &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;X&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;), &lt;span class=&quot;pl-k&quot;&gt;function&lt;/span&gt;(&lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;) &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[\[&lt;span class=&quot;pl-c1&quot;&gt;2&lt;/span&gt;\]\]))), &lt;span class=&quot;pl-v&quot;&gt;origin&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;1949-12-31&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  | which(&lt;span class=&quot;pl-smi&quot;&gt;ave.dates&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;1950-02-28&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; 59th doy&lt;/span&gt; |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;leap.ave&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; stack(&lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;\[\[&lt;span class=&quot;pl-c1&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;pl-c1&quot;&gt;59&lt;/span&gt;\]\], &lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;\[\[&lt;span class=&quot;pl-c1&quot;&gt;59&lt;/span&gt;\]\], &lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;\[\[&lt;span class=&quot;pl-c1&quot;&gt;60&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;pl-c1&quot;&gt;365&lt;/span&gt;\]\]) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; leap year version duplicates climatology from feb 28 for leap day&lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;leap.ave.bd&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; format(as\_date(&lt;span class=&quot;pl-c1&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;pl-c1&quot;&gt;366&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;origin&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;1951-12-31&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;), &lt;span class=&quot;pl-v&quot;&gt;format&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;%b-%d&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;clima&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; stack(&lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;leap.ave&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;ave&lt;/span&gt;) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; since starting with 1950, non-leap, non-leap, leap, non-leap, and then recycle. &lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;clima.fullsize&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; stack(replicate(&lt;span class=&quot;pl-c1&quot;&gt;17&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;clima&lt;/span&gt;)) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; 17 iterations of this 4 year sequence, now comparable in dimensions to anomalies object &lt;/span&gt; |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; as\_date(as.numeric(unlist(lapply(strsplit(names(&lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;), &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;span class=&quot;pl-cce&quot;&gt;\\\\&lt;/span&gt;.&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;), &lt;span class=&quot;pl-k&quot;&gt;function&lt;/span&gt;(&lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;) &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[\[&lt;span class=&quot;pl-c1&quot;&gt;2&lt;/span&gt;\]\]))), &lt;span class=&quot;pl-v&quot;&gt;origin&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;1949-12-31&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;years&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; unique(year(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;)) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;month.day&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; format(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;format&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;%b-%d&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — drop 2018, which is problematic because it&apos;s only a partial year&lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; subset(&lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;, which(&lt;span class=&quot;pl-k&quot;&gt;!&lt;/span&gt;year(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;) &lt;span class=&quot;pl-k&quot;&gt;%in%&lt;/span&gt; &lt;span class=&quot;pl-c1&quot;&gt;2018&lt;/span&gt;), &lt;span class=&quot;pl-v&quot;&gt;drop&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-c1&quot;&gt;T&lt;/span&gt;) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;\[&lt;span class=&quot;pl-k&quot;&gt;!&lt;/span&gt;year(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;) &lt;span class=&quot;pl-k&quot;&gt;%in%&lt;/span&gt; &lt;span class=&quot;pl-c1&quot;&gt;2018&lt;/span&gt;\] |
|  |  |
|  | print(length(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;) &lt;span class=&quot;pl-k&quot;&gt;==&lt;/span&gt; dim(&lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;)\[&lt;span class=&quot;pl-c1&quot;&gt;3&lt;/span&gt;\]) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;clima&lt;/span&gt; &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; recycles to length of anomalies &lt;/span&gt; |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — perform spot checks to ensure values are properly summed &lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;test.layers&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; floor(runif(&lt;span class=&quot;pl-c1&quot;&gt;20&lt;/span&gt;) &lt;span class=&quot;pl-k&quot;&gt;\*&lt;/span&gt; dim(&lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;)\[&lt;span class=&quot;pl-c1&quot;&gt;3&lt;/span&gt;\]) %&lt;span class=&quot;pl-k&quot;&gt;&amp;gt;&lt;/span&gt;% unique() |
|  | &lt;span class=&quot;pl-smi&quot;&gt;anomalies.test&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;\[\[&lt;span class=&quot;pl-smi&quot;&gt;test.layers&lt;/span&gt;\]\] |
|  | &lt;span class=&quot;pl-smi&quot;&gt;monthday.test&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;month.day&lt;/span&gt;\[&lt;span class=&quot;pl-smi&quot;&gt;test.layers&lt;/span&gt;\] |
|  | &lt;span class=&quot;pl-smi&quot;&gt;indices&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; unlist(lapply(&lt;span class=&quot;pl-smi&quot;&gt;monthday.test&lt;/span&gt;, &lt;span class=&quot;pl-k&quot;&gt;function&lt;/span&gt;(&lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;) which(&lt;span class=&quot;pl-smi&quot;&gt;leap.ave.bd&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;))) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; julian day values &lt;/span&gt; |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — these all check out — satisfied it&apos;s taking the correct sums — # &lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;sum.random&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;anomalies&lt;/span&gt;\[\[&lt;span class=&quot;pl-smi&quot;&gt;test.layers&lt;/span&gt;\]\] &lt;span class=&quot;pl-k&quot;&gt;+&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;leap.ave&lt;/span&gt;\[\[&lt;span class=&quot;pl-smi&quot;&gt;indices&lt;/span&gt;\]\] &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; manual combination of anomalies layer and climatology&lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;true.values&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt;\[\[&lt;span class=&quot;pl-smi&quot;&gt;test.layers&lt;/span&gt;\]\] &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; our combined temperature values layer &lt;/span&gt; |
|  | min(&lt;span class=&quot;pl-smi&quot;&gt;true.values&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;==&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;sum.random&lt;/span&gt;) |
|  |  |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; setZ(&lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;) |
|  | names(&lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt;) &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; as.character(&lt;span class=&quot;pl-smi&quot;&gt;dates&lt;/span&gt;) |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;temp&lt;/span&gt; |
|  | } |

&lt;/div&gt; &lt;/div&gt; &lt;/div&gt;&lt;/div&gt; &lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/acfb8181ebcdb5375c4d085064f22029/raw/3c166b63e2016e80412e3e440fa3222ef127dfed/clean_BEST_data.r) [ clean\_BEST\_data.r ](https://gist.github.com/a8dx/acfb8181ebcdb5375c4d085064f22029#file-clean_best_data-r) hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt; &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;
&lt;p&gt;You’ll notice some hard coded elements, such as the origin date and whether to drop partial years (since BEST is continuously updating their data, a portion of the current year will always be included in the most recent decadal extract). I’ve also included some tests to convince myself that the summed product mirrors values obtained from a manually searched climatology layer combined with the anomalies layer for a specified date.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>andagostino</name>
          
          
        </author>
      

      
        <category term="climate change" />
      
        <category term="Data" />
      
        <category term="R" />
      

      

      
        <summary type="html">You may have recently seen air quality maps produced by the Berkeley Earth group, especially in the wake of the horrific Camp Fire whose death toll now exceeds 80. For example, here’s their real-time visualization of PM2.5 concentrations.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Bounding Boxes for All US States</title>
      
      
      <link href="https://anthonylouisdagostino.com/bounding-boxes-for-all-us-states/" rel="alternate" type="text/html" title="Bounding Boxes for All US States" />
      
      <published>2018-11-23T09:13:00+00:00</published>
      <updated>2018-11-23T09:13:00+00:00</updated>
      <id>https://anthonylouisdagostino.com/bounding-boxes-for-all-us-states</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/bounding-boxes-for-all-us-states/">&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/11/boundingboxes.png?resize=712%2C534&amp;amp;ssl=1&quot; alt=&quot;boundingboxes&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Sometimes you come across an API that requires bounding box coordinates to subset your query, but which doesn’t offer an interactive map to actually create one and extract those min/max values. No fear, here are the extents of each US state and territory in NAD83 coordinates using the &lt;a href=&quot;https://www.census.gov/geo/maps-data/data/cbf/cbf_state.html&quot;&gt;2017 US Census 1:500,000 shapefile&lt;/a&gt;.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist93176793&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-us_state_bounding_boxes-csv&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-csv  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;blob-interaction-bar&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-search&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M11.5 7a4.499 4.499 0 11-8.998 0A4.499 4.499 0 0111.5 7zm-.82 4.74a6 6 0 111.06-1.06l3.04 3.04a.75.75 0 11-1.06 1.06l-3.04-3.04z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
 &lt;input aria-label=&quot;Search this file…&quot; autocapitalize=&quot;off&quot; class=&quot;form-control js-csv-filter-field blob-filter&quot; name=&quot;filter&quot; placeholder=&quot;Search this file…&quot; type=&quot;text&quot; /&gt;&amp;lt;/input&amp;gt;&lt;/div&gt;&lt;div class=&quot;markdown-body js-check-bidi&quot; data-hpc=&quot;&quot; data-line-alert=&quot;before&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;&lt;/template&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;|  |  | STATEFP | STUSPS | NAME | xmin | ymin | xmax | ymax |
|---|---|---|---|---|---|---|---|---|
|  | 1 | 01 | AL | Alabama | -88.473227 | 30.223334 | -84.88908 | 35.008028 |
|  | 2 | 02 | AK | Alaska | -179.148909 | 51.214183 | 179.77847 | 71.365162 |
|  | 3 | 60 | AS | American Samoa | -171.089874 | -14.548699 | -168.1433 | -11.046934 |
|  | 4 | 04 | AZ | Arizona | -114.81651 | 31.332177 | -109.045223 | 37.00426 |
|  | 5 | 05 | AR | Arkansas | -94.617919 | 33.004106 | -89.644395 | 36.4996 |
|  | 6 | 06 | CA | California | -124.409591 | 32.534156 | -114.131211 | 42.009518 |
|  | 7 | 08 | CO | Colorado | -109.060253 | 36.992426 | -102.041524 | 41.003444 |
|  | 8 | 69 | MP | Commonwealth of the Northern Mariana Islands | 144.886331 | 14.110472 | 146.064818 | 20.553802 |
|  | 9 | 09 | CT | Connecticut | -73.727775 | 40.980144 | -71.786994 | 42.050587 |
|  | 10 | 10 | DE | Delaware | -75.788658 | 38.451013 | -75.048939 | 39.839007 |
|  | 11 | 11 | DC | District of Columbia | -77.119759 | 38.791645 | -76.909395 | 38.99511 |
|  | 12 | 12 | FL | Florida | -87.634938 | 24.523096 | -80.031362 | 31.000888 |
|  | 13 | 13 | GA | Georgia | -85.605165 | 30.357851 | -80.839729 | 35.000659 |
|  | 14 | 66 | GU | Guam | 144.618068 | 13.234189 | 144.956712 | 13.654383 |
|  | 15 | 15 | HI | Hawaii | -178.334698 | 18.910361 | -154.806773 | 28.402123 |
|  | 16 | 16 | ID | Idaho | -117.243027 | 41.988057 | -111.043564 | 49.001146 |
|  | 17 | 17 | IL | Illinois | -91.513079 | 36.970298 | -87.494756 | 42.508481 |
|  | 18 | 18 | IN | Indiana | -88.09776 | 37.771742 | -84.784579 | 41.760592 |
|  | 19 | 19 | IA | Iowa | -96.639704 | 40.375501 | -90.140061 | 43.501196 |
|  | 20 | 20 | KS | Kansas | -102.051744 | 36.993016 | -94.588413 | 40.003162 |
|  | 21 | 21 | KY | Kentucky | -89.571509 | 36.497129 | -81.964971 | 39.147458 |
|  | 22 | 22 | LA | Louisiana | -94.043147 | 28.928609 | -88.817017 | 33.019457 |
|  | 23 | 23 | ME | Maine | -71.083924 | 42.977764 | -66.949895 | 47.459686 |
|  | 24 | 24 | MD | Maryland | -79.487651 | 37.911717 | -75.048939 | 39.723043 |
|  | 25 | 25 | MA | Massachusetts | -73.508142 | 41.237964 | -69.928393 | 42.886589 |
|  | 26 | 26 | MI | Michigan | -90.418136 | 41.696118 | -82.413474 | 48.2388 |
|  | 27 | 27 | MN | Minnesota | -97.239209 | 43.499356 | -89.491739 | 49.384358 |
|  | 28 | 28 | MS | Mississippi | -91.655009 | 30.173943 | -88.097888 | 34.996052 |
|  | 29 | 29 | MO | Missouri | -95.774704 | 35.995683 | -89.098843 | 40.61364 |
|  | 30 | 30 | MT | Montana | -116.050003 | 44.358221 | -104.039138 | 49.00139 |
|  | 31 | 31 | NE | Nebraska | -104.053514 | 39.999998 | -95.30829 | 43.001708 |
|  | 32 | 32 | NV | Nevada | -120.005746 | 35.001857 | -114.039648 | 42.002207 |
|  | 33 | 33 | NH | New Hampshire | -72.557247 | 42.69699 | -70.610621 | 45.305476 |
|  | 34 | 34 | NJ | New Jersey | -75.559614 | 38.928519 | -73.893979 | 41.357423 |
|  | 35 | 35 | NM | New Mexico | -109.050173 | 31.332301 | -103.001964 | 37.000232 |
|  | 36 | 36 | NY | New York | -79.762152 | 40.496103 | -71.856214 | 45.01585 |
|  | 37 | 37 | NC | North Carolina | -84.321869 | 33.842316 | -75.460621 | 36.588117 |
|  | 38 | 38 | ND | North Dakota | -104.0489 | 45.935054 | -96.554507 | 49.000574 |
|  | 39 | 39 | OH | Ohio | -84.820159 | 38.403202 | -80.518693 | 41.977523 |
|  | 40 | 40 | OK | Oklahoma | -103.002565 | 33.615833 | -94.430662 | 37.002206 |
|  | 41 | 41 | OR | Oregon | -124.566244 | 41.991794 | -116.463504 | 46.292035 |
|  | 42 | 42 | PA | Pennsylvania | -80.519891 | 39.7198 | -74.689516 | 42.26986 |
|  | 43 | 72 | PR | Puerto Rico | -67.945404 | 17.88328 | -65.220703 | 18.515683 |
|  | 44 | 44 | RI | Rhode Island | -71.862772 | 41.146339 | -71.12057 | 42.018798 |
|  | 45 | 45 | SC | South Carolina | -83.35391 | 32.0346 | -78.54203 | 35.215402 |
|  | 46 | 46 | SD | South Dakota | -104.057698 | 42.479635 | -96.436589 | 45.94545 |
|  | 47 | 47 | TN | Tennessee | -90.310298 | 34.982972 | -81.6469 | 36.678118 |
|  | 48 | 48 | TX | Texas | -106.645646 | 25.837377 | -93.508292 | 36.500704 |
|  | 49 | 78 | VI | United States Virgin Islands | -65.085452 | 17.673976 | -64.564907 | 18.412655 |
|  | 50 | 49 | UT | Utah | -114.052962 | 36.997968 | -109.041058 | 42.001567 |
|  | 51 | 50 | VT | Vermont | -73.43774 | 42.726853 | -71.464555 | 45.016659 |
|  | 52 | 51 | VA | Virginia | -83.675395 | 36.540738 | -75.242266 | 39.466012 |
|  | 53 | 53 | WA | Washington | -124.763068 | 45.543541 | -116.915989 | 49.002494 |
|  | 54 | 54 | WV | West Virginia | -82.644739 | 37.201483 | -77.719519 | 40.638801 |
|  | 55 | 55 | WI | Wisconsin | -92.888114 | 42.491983 | -86.805415 | 47.080621 |
|  | 56 | 56 | WY | Wyoming | -111.056888 | 40.994746 | -104.05216 | 45.005904 |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/2340f9527af64f8ef8439366de981168/raw/81d876daea10eab5c2675811c39bcd18a79a9212/US_State_Bounding_Boxes.csv)  
 [  
 US\_State\_Bounding\_Boxes.csv  
 ](https://gist.github.com/a8dx/2340f9527af64f8ef8439366de981168#file-us_state_bounding_boxes-csv)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;To back these out, I use the incredible &lt;a href=&quot;https://r-spatial.github.io/sf/&quot;&gt;simple features&lt;/a&gt; R package written by Edzer Pebesma.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="Geography" />
      
        <category term="R" />
      

      

      
        <summary type="html"></summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">A Better ZIP5-County Crosswalk</title>
      
      
      <link href="https://anthonylouisdagostino.com/a-better-zip5-county-crosswalk/" rel="alternate" type="text/html" title="A Better ZIP5-County Crosswalk" />
      
      <published>2018-03-28T19:55:09+00:00</published>
      <updated>2018-03-28T19:55:09+00:00</updated>
      <id>https://anthonylouisdagostino.com/a-better-zip5-county-crosswalk</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/a-better-zip5-county-crosswalk/">&lt;p&gt;I use a healthcare expenditure dataset with observations geographically coded at the 5-digit zipcode level, but I’d also like to know which county an observation ‘belongs’ to. Maybe I want to cluster standard errors by county, or control for county-specific trends. You’d imagine this would be straightforward, but I haven’t yet found a government crosswalk that is comprehensive in all the ZIP5s that appear in my data. What follows is the best solution I’m aware of, to ensure that I match as many ZIP5s as possible. While this only increases the number of ZIP5-county matches by about 110 over what HUD offers, it’s an improvement of more than 6,000 over the Census crosswalk.&lt;/p&gt;

&lt;p&gt;Mapping ZIP5s to counties is slightly tricky because they,&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Refer to postal routes, and are therefore collections of lines, not polygons&lt;/li&gt;
  &lt;li&gt;May straddle multiple counties, necessitating an apportionment decision&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href=&quot;https://www.census.gov/geo/reference/zctas.html&quot;&gt;US Census Bureau&lt;/a&gt; addresses 1.) by devising a Zip Code Tabulation Area (ZCTA) which approximates the area served in a ZIP using Census blocks. These are by and large mapped 1:1 such that, for example, the 31211 zip5 in Macon, Georgia also appears as 31211 in a ZCTA database.&lt;/p&gt;

&lt;p&gt;As for 2.), Available crosswalks do provide several importance measures to aid the apportionment decision, under the assumption that your analysis requires that a singular county be mapped to a ZIP. HUD provides &lt;a href=&quot;https://www.huduser.gov/portal/datasets/usps_crosswalk.html&quot;&gt;crosswalks updated quarterly&lt;/a&gt; for several administrative levels (CBSA, CBSA division, Census tract, etc.), and &lt;a href=&quot;https://www.huduser.gov/portal/datasets/usps_crosswalk.html#codebook&quot;&gt;includes the share&lt;/a&gt; of that zipcode’s residential addresses, business addresses, other addresses, and total addresses that lie in any of the relevant counties.&lt;/p&gt;

&lt;p&gt;So we’re done, right? Not yet. First, it’s not clear that an address count is ideal. Imagine a scenario where a zipcode broaches several adjacent counties and features two types of housing: senior housing and homes occupied by extended families. If the average number of residents per unit in the retirement homes is 1.5, but 6 in the latter, then the share of addresses is a very imperfect proxy for population shares.&lt;/p&gt;

&lt;p&gt;The &lt;a href=&quot;https://www2.census.gov/geo/docs/maps-data/data/rel/zcta_county_rel_10.txt&quot;&gt;Census Bureau&lt;/a&gt; addresses this with their lookup table and includes 2010 Census population data as well as area (total and land-only). &lt;a href=&quot;https://www2.census.gov/geo/pdfs/maps-data/data/rel/explanation_zcta_county_rel_10.pdf&quot;&gt;Here is a helpful guide&lt;/a&gt; they’ve created. The crosswalk includes some relevant information in the county –&amp;gt; ZIP5 direction as well, such as — &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;COPOPPCT&lt;/code&gt;— the county population percentage residing in that ZIP5.&lt;/p&gt;

&lt;p&gt;Let’s first get the &lt;a href=&quot;https://www.census.gov/geo/reference/codes/cou.html&quot;&gt;Census listing of county FIPS codes and names&lt;/a&gt;.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist88663164&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-gistfile1-txt&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-text  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | clear |
|---|---|
|  | clear all |
|  | clear matrix |
|  | set more off |
|  | set maxvar 25000, permanently |
|  |  |
|  |  |
|  |  |
|  | loc basePath &quot;&amp;lt;your path here&amp;gt;&quot; |
|  |  |
|  |  |
|  | // read in county codes, downloaded from: https://www.census.gov/geo/reference/codes/cou.html |
|  | import delimited using &quot;`basePath&apos;/national\_county.txt&quot;, delim(&quot;,&quot;) clear stringcols(\_all) varnames(nonames) |
|  |  |
|  | ren v1 state |
|  | ren v2 state\_fips |
|  | ren v3 county\_fips |
|  | ren v4 county\_name |
|  | ren v5 fipsclasscode |
|  |  |
|  | gen county = state\_fips + county\_fips |
|  |  |
|  | tempfile counties |
|  | save `counties&apos; |
|  |  |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/20a27a330a14ba3907fff455e5994fed/raw/e80d962fb3e7efa5574c19de1e053b862521f2bb/gistfile1.txt)  
 [  
 gistfile1.txt  
 ](https://gist.github.com/a8dx/20a27a330a14ba3907fff455e5994fed#file-gistfile1-txt)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;The &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;county&lt;/code&gt; variable is concatenated from the state FIPS and county FIPS codes, and will be used to link with the ZIP-County crosswalks.&lt;/p&gt;

&lt;p&gt;Now let’s start with that &lt;a href=&quot;https://www.huduser.gov/portal/datasets/usps_crosswalk.html&quot;&gt;HUD crosswalk&lt;/a&gt;. I’ve randomly selected the December 2017 version, but the following applies to any vintage. This file includes 39,455 unique zipcodes. There is substantial change over time, as an equally-randomly selected March 2011 file features 36,413 unique zipcodes. We’ll merge this in with our county FIPS file and identify the county for which a given zipcode has the largest share of total addresses, and residential addresses in.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist88663626&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-gistfile1-txt&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-text  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | import excel using &quot;`basePath&apos;/ZIP\_COUNTY\_122017.xlsx&quot;, firstrow clear |
|---|---|
|  |  |
|  | // 39455 unique entries before any transformations |
|  | merge m:1 county using `counties&apos;, update replace |
|  | tab \_merge |
|  | drop if \_merge == 2 |
|  | drop \_merge |
|  |  |
|  | gsort zip -tot\_ratio |
|  | bys zip: gen totOrder = \_n |
|  |  |
|  | gsort zip -res\_ratio |
|  | bys zip: gen resOrder = \_n |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/e739c69c435911fb76a3a77004782468/raw/76ef1d4601f6474a389f27ecf4c39d40ba08c23f/gistfile1.txt)  
 [  
 gistfile1.txt  
 ](https://gist.github.com/a8dx/e739c69c435911fb76a3a77004782468#file-gistfile1-txt)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/03/screen-shot-2018-03-28-at-11-52-48-am.png?resize=712%2C233&amp;amp;ssl=1&quot; alt=&quot;Screen Shot 2018-03-28 at 11.52.48 AM&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The vast majority of county matches with the greatest share of residential addresses also have the highest share of total addresses. We’ll opt for the former as our apportioning variable. And to give you a sense of how important apportionment is, more than 10,000 zip5’s reside in at least two counties.&lt;/p&gt;

&lt;p&gt;Is this problematic in any way? Let’s first inspect the residential address share for those counties that were first in the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gsort&lt;/code&gt; sorting above. As expected, the vast majority of them are above 50%, with the lion’s share at 100%. These “100% zipcodes” are interior to a single county, at least in regards to residential address location. &lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/03/resordershare.png?resize=712%2C518&amp;amp;ssl=1&quot; alt=&quot;ResOrderShare.png&quot; /&gt;&lt;/p&gt;

&lt;p&gt;What’s troubling is that mass at 0%. The sort function should not have loaded on these counties at all, yet eye-balling them suggests they are exclusively contained within single counties, so there’s not a problem. In fact, these zipcodes may simply contain only business/other addresses, in which case then they’re unlikely to further our healthcare expenditure data analysis.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist88664098&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-gistfile1-txt&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-text  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | import excel using &quot;`basePath&apos;/ZIP\_COUNTY\_122017.xlsx&quot;, firstrow clear |
|---|---|
|  |  |
|  | // 39455 unique entries before any transformations |
|  | merge m:1 county using `counties&apos;, update replace |
|  | tab \_merge |
|  | drop if \_merge == 2 |
|  | drop \_merge |
|  |  |
|  | gsort zip -tot\_ratio |
|  | bys zip: gen totOrder = \_n |
|  |  |
|  | gsort zip -res\_ratio |
|  | bys zip: gen resOrder = \_n |
|  |  |
|  | tw (hist res\_ratio if resOrder == 1), xtitle(&quot;Residential Address Percent&quot;) graphregion(color(white) lwidth(large)) |
|  | keep if resOrder == 1 |
|  |  |
|  | // some basic cleaning |
|  | replace county\_name = &quot;Oglala Lakota, SD&quot; if county == &quot;46102&quot; |
|  | replace state = &quot;SD&quot; if county == &quot;46102&quot; |
|  | replace state\_fips = &quot;46&quot; if county == &quot;46102&quot; |
|  | replace county\_fips = &quot;102&quot; if county == &quot;46102&quot; |
|  |  |
|  | // see http://www.nws.noaa.gov/om/notification/scn17-57kusilvak\_ak.htm |
|  | replace county\_name = &quot;Kusilvak Census Area&quot; if county == &quot;02158&quot; |
|  | replace state = &quot;AK&quot; if county == &quot;02158&quot; |
|  | replace state\_fips = &quot;02&quot; if county == &quot;02158&quot; |
|  | replace county\_fips = &quot;158&quot; if county == &quot;02158&quot; |
|  |  |
|  | drop resOrder totOrder |
|  | ren zip zip5 |
|  | destring, replace |
|  |  |
|  | tempfile zip5\_county |
|  | save `zip5\_county&apos; |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/65d82728d11c3fa5f6c5f10ba197eb72/raw/6c19a1209b03bd9a915335a9c5261a283069bddf/gistfile1.txt)  
 [  
 gistfile1.txt  
 ](https://gist.github.com/a8dx/65d82728d11c3fa5f6c5f10ba197eb72#file-gistfile1-txt)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;We’ll now download that &lt;a href=&quot;https://www2.census.gov/geo/docs/maps-data/data/rel/zcta_county_rel_10.txt&quot;&gt;Census crosswalk,&lt;/a&gt; and perform some similar procedures, again apportioning to the county for which the largest share of a zipcode’s population resides.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist88664158&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-gistfile1-txt&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-text  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | // downloaded from: https://www2.census.gov/geo/docs/maps-data/data/rel/zcta\_county\_rel\_10.txt |
|---|---|
|  | import delimited using &quot;`basePath&apos;/zcta\_county\_rel\_10.txt&quot;, delim(&quot;,&quot;) clear stringcols(\_all) |
|  |  |
|  | \*\* 33120 unique ZCTA5 |
|  | ren state state\_fips\_rel |
|  | ren county county\_fips\_rel |
|  |  |
|  | gen county = state\_fips\_rel + county\_fips\_rel |
|  |  |
|  | merge m:1 county using `counties&apos;, update replace |
|  | tab \_merge |
|  | drop if \_merge == 2 |
|  | drop \_merge |
|  |  |
|  | destring, replace |
|  |  |
|  | ren state state\_rel |
|  | ren county county\_rel |
|  | ren county\_name county\_name\_rel |
|  |  |
|  | // apportionment: keep county matches with largest pop ZCTA5 share |
|  | gsort zcta5 -poppt |
|  | bys zcta5: gen popOrder = \_n |
|  | keep if popOrder == 1 |
|  |  |
|  | ren zpop pop\_zip5 |
|  | ren zcta5 zip5 |
|  | keep zip5 pop\_zip5 zpoppct county\_rel state\_rel county\_name\_rel state\_fips\_rel county\_fips\_rel |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/41a46d3ca4e9d24659c7f7fda9d6e923/raw/9aec8477d87fbf51f99c96b3bd1b11fc719c3ab2/gistfile1.txt)  
 [  
 gistfile1.txt  
 ](https://gist.github.com/a8dx/41a46d3ca4e9d24659c7f7fda9d6e923#file-gistfile1-txt)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;I’ve suffixed several of the variables with “_rel” since a naive merge with our HUD crosswalk generates several merge conflicts. We can more carefully identify disagreements between the two crosswalks this way.&lt;/p&gt;

&lt;p&gt;We now merge and get the following:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/03/screen-shot-2018-03-28-at-12-22-40-pm.png?resize=554%2C468&amp;amp;ssl=1&quot; alt=&quot;Screen Shot 2018-03-28 at 12.22.40 PM&quot; /&gt;&lt;/p&gt;

&lt;p&gt;so that there’s 33,014 zipcodes appearing in both crosswalks, while the HUD database adds 106 to those found in the Census file.&lt;/p&gt;

&lt;p&gt;Here’s our first problem – major inconsistencies in the county results from our two apportionment procedures. &lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/03/screen-shot-2018-03-28-at-12-26-58-pm.png?resize=657%2C425&amp;amp;ssl=1&quot; alt=&quot;Screen Shot 2018-03-28 at 12.26.58 PM&quot; /&gt;&lt;/p&gt;

&lt;p&gt;That’s pretty bad, but it’s even worse when we’re not even agreeing on the same state.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/03/screen-shot-2018-03-28-at-12-29-08-pm.png?resize=712%2C76&amp;amp;ssl=1&quot; alt=&quot;Screen Shot 2018-03-28 at 12.29.08 PM.png&quot; /&gt;&lt;/p&gt;

&lt;p&gt;In these cases, HUD tells us that 100% of residential addresses are located in the designated county, while Census gives values for these 3 of between 53% and 69%. Population apportionment would therefore direct us to the state_rel values and respective counties.&lt;/p&gt;

&lt;p&gt;So a final decision to make is how these disagreements should be reconciled. The good news is that this need applies only to 357 zipcodes which have conflicting information coming from both crosswalks. In many cases, county information is only coming from one of the crosswalks.&lt;/p&gt;

&lt;p&gt;I think the optimal approach is the following and is the algorithm used in producing the following dataset:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Rely first on Census population-apportioned county matching&lt;/li&gt;
  &lt;li&gt;Then fill missing values with HUD’s residential address apportioned matches&lt;/li&gt;
  &lt;li&gt;When both are present but in conflict, rely on Census values which should be considered to have more integrity than HUD values derived from Census raw data&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The final dataset will have 39,561 unique zipcodes, if you were to download the HUD crosswalk vintage referred to above. It marries the reliability of Census population data without the potential for errors percolating from another government agency’s analysis, with the timeliness of the quarterly HUD datasets which expands the scope of included ZIP5s.&lt;/p&gt;

&lt;p&gt;You can &lt;a href=&quot;https://www.dropbox.com/s/mycgewaljt9ic17/ZIP5_County_Crosswalk.dta?dl=0&quot;&gt;download the final Stata .dta crosswalk file here&lt;/a&gt;, while the &lt;a href=&quot;https://gist.github.com/a8dx/7e9d5af24101fc66aafa739577713b59&quot;&gt;entire Stata script is available here&lt;/a&gt;.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>andagostino</name>
          
          
        </author>
      

      
        <category term="Data" />
      
        <category term="Geography" />
      

      

      
        <summary type="html">I use a healthcare expenditure dataset with observations geographically coded at the 5-digit zipcode level, but I’d also like to know which county an observation ‘belongs’ to. Maybe I want to cluster standard errors by county, or control for county-specific trends. You’d imagine this would be straightforward, but I haven’t yet found a government crosswalk that is comprehensive in all the ZIP5s that appear in my data. What follows is the best solution I’m aware of, to ensure that I match as many ZIP5s as possible. While this only increases the number of ZIP5-county matches by about 110 over what HUD offers, it’s an improvement of more than 6,000 over the Census crosswalk.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Large Stata Datasets and False Errors about ‘Duplicates’</title>
      
      
      <link href="https://anthonylouisdagostino.com/large-stata-datasets-and-false-errors-about-duplicates/" rel="alternate" type="text/html" title="Large Stata Datasets and False Errors about &apos;Duplicates&apos;" />
      
      <published>2018-02-01T01:49:42+00:00</published>
      <updated>2018-02-01T01:49:42+00:00</updated>
      <id>https://anthonylouisdagostino.com/large-stata-datasets-and-false-errors-about-duplicates</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/large-stata-datasets-and-false-errors-about-duplicates/">&lt;p&gt;Variable storage types exercise more importance when working with larger datasets, and variables with more digits. I’m reminded of this because of an error message Stata threw while trying to perform a long &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;reshape&lt;/code&gt;, claiming duplicate entries of the ID variable. That was obviously not the case, since the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;_n&lt;/code&gt; id was uniquely created, and the value of each visibly corresponded to its row index.&lt;/p&gt;

&lt;p&gt;The problem is Stata’s default type for numeric is as a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;float&lt;/code&gt;. Under many circumstances, that’s fine for either integers or decimal numeric objects. But with N≥ 20 million, the dataset that prompted the error is butting up against precision limits since a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;float&lt;/code&gt; &lt;a href=&quot;https://www.stata.com/manuals13/ddatatypes.pdf&quot;&gt;is accurate to 7 digits.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The solution is to use a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;double&lt;/code&gt; type instead, which can reliably hold up to 16 digits. Anything larger, you’re likely best off working with strings.&lt;/p&gt;

&lt;p&gt;Specifying the storage type is straightforward, like so:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gen double id = _n&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So when you’re trying to reshape a large dataset and Stata quits, even though you &lt;em&gt;know&lt;/em&gt; you’ve satisfied uniqueness in your identifying variable, &lt;strong&gt;double&lt;/strong&gt;-check your ID’s storage type.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>andagostino</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      

      

      
        <summary type="html">Variable storage types exercise more importance when working with larger datasets, and variables with more digits. I’m reminded of this because of an error message Stata threw while trying to perform a long reshape, claiming duplicate entries of the ID variable. That was obviously not the case, since the _n id was uniquely created, and the value of each visibly corresponded to its row index.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Identify nth-Degree Neighbors Using R’s Simple Features Package, Simply</title>
      
      
      <link href="https://anthonylouisdagostino.com/identify-nth-degree-neighbors-using-rs-simple-features-package-simply/" rel="alternate" type="text/html" title="Identify nth-Degree Neighbors Using R&apos;s Simple Features Package, Simply" />
      
      <published>2018-01-19T21:21:32+00:00</published>
      <updated>2018-01-19T21:21:32+00:00</updated>
      <id>https://anthonylouisdagostino.com/identify-nth-degree-neighbors-using-rs-simple-features-package-simply</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/identify-nth-degree-neighbors-using-rs-simple-features-package-simply/">&lt;p&gt;You’re more likely to complain about the &lt;a href=&quot;https://www.youtube.com/watch?v=4IRB0sxw-YU&quot;&gt;neighbors upstairs who are making noise&lt;/a&gt; after midnight than those in an apartment two buildings away. Proximity matters and that’s patently obvious, but oftentimes it takes a bit of work to identify who is close and who isn’t. While raster data is packaged in a consistent gridded format for which inverse distance weighting schemes readily can be applied, shapefiles with oddly-shaped features, like &lt;a href=&quot;https://www.washingtonpost.com/news/wonk/wp/2014/05/15/americas-most-gerrymandered-congressional-districts/?utm_term=.c9373f790cfe&quot;&gt;these gerrymandered districts&lt;/a&gt;, may present more of a challenge. Fortunately the &lt;a href=&quot;https://github.com/r-spatial/sf&quot;&gt;simple features library in R&lt;/a&gt; can save the day and with little sweat on your brow.&lt;/p&gt;

&lt;p&gt;This brief tutorial walks through an example of identifying all first- (all my neighbors who share a common edge with me) and second-degree neighbors (all neighbors of my neighbors) of a given county in North Carolina. Since the map is bundled with the library install, no additional downloading is needed. Everything is generalizable and extending to higher-order neighbors is straightforward, but probably unnecessary. While some great R packages exist to perform this computation in a graph setting, I think sf is superior for spatial data.&lt;/p&gt;

&lt;p&gt;A couple years ago my buddy and co-author &lt;a href=&quot;http://www.eyalfrank.com/&quot;&gt;Eyal Frank&lt;/a&gt; published a post outlining a similar ArcGIS workflow, but since then the &lt;a href=&quot;https://github.com/r-spatial/sf&quot;&gt;sf library&lt;/a&gt; for R has come online and greatly expands R’s spatial analysis capabilities, to the extent that much of what was previously achievable (and easy) through Arc products can now be run natively in R – plus you don’t have to pray that your session will complete before the program crashes. We can achieve similar ends with less code, and without giving ESRI all of our money. &lt;em&gt;You can instead give it to &lt;a href=&quot;https://www.r-consortium.org/projects&quot;&gt;R Consortium&lt;/a&gt; which provides financial support to awesome R developers like &lt;a href=&quot;https://github.com/edzer&quot;&gt;Edzer Pebesma&lt;/a&gt; and &lt;a href=&quot;https://github.com/jeroen&quot;&gt;Jeroen Ooms&lt;/a&gt; for new creations.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The complete code is &lt;a href=&quot;https://gist.github.com/a8dx/7f588b7da531e93049b2b269a3670c89&quot;&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Let’s start with the North Carolina county-level map packaged with your sf install, as an example. Let’s select Chatham County, since it’s in the middle of the state and will therefore give us many second-degree neighbors without needing to cross state lines. For now we start with the entire state, and then will drill down to just Chatham.&lt;/p&gt;

&lt;p&gt;The heavy lifting is done in the FIRSTdegreeNeighbors function which uses &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;st_touches&lt;/code&gt; to return index values corresponding to all adjacent entries. This is performed for every county in the state and with some finagling, we can produce a data frame of all positive matches.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist85730378&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-firstdegreeneighbors-r&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-r  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | &lt;span class=&quot;pl-en&quot;&gt;FIRSTdegreeNeighbors&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;function&lt;/span&gt;(&lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;) { |
|---|---|
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — use sf functionality and output to sparse matrix format for minimizing footprint &lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;first.neighbor&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; st\_touches(&lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;sparse&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-c1&quot;&gt;TRUE&lt;/span&gt;) |
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — convert results to data.frame (via sparse matrix)&lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;n.ids&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; sapply(&lt;span class=&quot;pl-smi&quot;&gt;first.neighbor&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;length&lt;/span&gt;) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;vals&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; unlist(&lt;span class=&quot;pl-smi&quot;&gt;first.neighbor&lt;/span&gt;) |
|  | &lt;span class=&quot;pl-smi&quot;&gt;out&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; sparseMatrix(&lt;span class=&quot;pl-smi&quot;&gt;vals&lt;/span&gt;, rep(seq\_along(&lt;span class=&quot;pl-smi&quot;&gt;n.ids&lt;/span&gt;), &lt;span class=&quot;pl-smi&quot;&gt;n.ids&lt;/span&gt;)) |
|  |  |
|  | &lt;span class=&quot;pl-smi&quot;&gt;out.summ&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; summary(&lt;span class=&quot;pl-smi&quot;&gt;out&lt;/span&gt;) &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — this is currently only generating row values, need to map to actual obs \[next line\]&lt;/span&gt; |
|  | &lt;span class=&quot;pl-k&quot;&gt;data.frame&lt;/span&gt;(&lt;span class=&quot;pl-v&quot;&gt;county&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[&lt;span class=&quot;pl-smi&quot;&gt;out.summ&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;j&lt;/span&gt;,\]&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;NAME&lt;/span&gt;, |
|  | &lt;span class=&quot;pl-v&quot;&gt;countyid&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[&lt;span class=&quot;pl-smi&quot;&gt;out.summ&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;j&lt;/span&gt;,\]&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;CNTY\_ID&lt;/span&gt;, |
|  | &lt;span class=&quot;pl-v&quot;&gt;firstdegreeneighbors&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[&lt;span class=&quot;pl-smi&quot;&gt;out.summ&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;i&lt;/span&gt;,\]&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;NAME&lt;/span&gt;, |
|  | &lt;span class=&quot;pl-v&quot;&gt;firstdegreeneighborid&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;x&lt;/span&gt;\[&lt;span class=&quot;pl-smi&quot;&gt;out.summ&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;i&lt;/span&gt;,\]&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;CNTY\_ID&lt;/span&gt;, |
|  | &lt;span class=&quot;pl-v&quot;&gt;stringsAsFactors&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-c1&quot;&gt;FALSE&lt;/span&gt;) |
|  | } |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/318ad0cb61befb3a740907e28792a6cd/raw/23b4e0fac62097a0c821da31f71c4e2219a5b90e/FIRSTdegreeNeighbors.r)  
 [  
 FIRSTdegreeNeighbors.r  
 ](https://gist.github.com/a8dx/318ad0cb61befb3a740907e28792a6cd#file-firstdegreeneighbors-r)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;Since we’re relying on sparse matrices to hold our results, you could supply a large feature collection (e.g., 50,000+) and not worry about crashing R because you’ve exceeded memory caps that dense matrices will blow through. [h/t &lt;a href=&quot;https://stackoverflow.com/questions/4942361/how-to-turn-a-list-of-lists-to-a-sparse-matrix-in-r-without-using-lapply?rq=1&quot;&gt;Aaron&lt;/a&gt;]&lt;/p&gt;

&lt;p&gt;The output is a symmetric matrix which can be plotted using the county ID values, where count is the number of first degree neighbor matches over some defined hexagonal size. You can control this size, which directly affects total counts per ‘bin.’ You might use this as a sanity check that only nearby features are getting picked up, but why stop there — we should plot the results and make sure it’s doing what we expect.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/01/nc_firstneighbors_symmetricmatrix.png?resize=594%2C522&amp;amp;ssl=1&quot; alt=&quot;NC_FirstNeighbors_SymmetricMatrix&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The code shows how easy it is to iterate the process through subsetting and merging to generate the set of second-degree neighbors. There’s some extra work to ensure that none of our first-degree neighbors appear in the list of second-degree neighbors, but that’s easy enough to do using set functions and subsets.&lt;/p&gt;

&lt;style&gt;.gist table { margin-bottom: 0; }&lt;/style&gt;
&lt;div class=&quot;gist&quot; id=&quot;gist85730804&quot; style=&quot;tab-size: 8&quot;&gt;&lt;div class=&quot;gist-file&quot; translate=&quot;no&quot;&gt;&lt;div class=&quot;gist-data&quot;&gt;&lt;div class=&quot;js-gist-file-update-container js-task-list-container file-box&quot;&gt;&lt;div class=&quot;file my-2&quot; id=&quot;file-seconddegreeneighbors-r&quot;&gt;&lt;div class=&quot;Box-body p-0 blob-wrapper data type-r  &quot; itemprop=&quot;text&quot;&gt;&lt;div class=&quot;js-check-bidi js-blob-code-container blob-code-content&quot;&gt; &lt;template class=&quot;js-file-alert-template&quot;&gt;&lt;/template&gt;

&lt;div class=&quot;flash flash-warn flash-full d-flex flex-items-center&quot; data-view-component=&quot;true&quot;&gt; &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt; &lt;span&gt;  
 This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.  
 [Learn more about bidirectional Unicode characters](https://github.co/hiddenchars)  
 &lt;/span&gt;

&lt;div class=&quot;flash-action&quot; data-view-component=&quot;true&quot;&gt; [ Show hidden characters  ](&amp;lt;&amp;gt;)&lt;/div&gt;&lt;/div&gt;  
&lt;template class=&quot;js-line-alert-template&quot;&gt;  
 &lt;span aria-label=&quot;This line has hidden Unicode characters&quot; class=&quot;line-alert tooltipped tooltipped-e&quot; data-view-component=&quot;true&quot;&gt;  
 &lt;svg aria-hidden=&quot;true&quot; class=&quot;octicon octicon-alert&quot; data-view-component=&quot;true&quot; height=&quot;16&quot; version=&quot;1.1&quot; viewbox=&quot;0 0 16 16&quot; width=&quot;16&quot;&gt; &lt;path d=&quot;M8.22 1.754a.25.25 0 00-.44 0L1.698 13.132a.25.25 0 00.22.368h12.164a.25.25 0 00.22-.368L8.22 1.754zm-1.763-.707c.659-1.234 2.427-1.234 3.086 0l6.082 11.378A1.75 1.75 0 0114.082 15H1.918a1.75 1.75 0 01-1.543-2.575L6.457 1.047zM9 11a1 1 0 11-2 0 1 1 0 012 0zm-.25-5.25a.75.75 0 00-1.5 0v2.5a.75.75 0 001.5 0v-2.5z&quot; fill-rule=&quot;evenodd&quot;&gt;&lt;/path&gt;&lt;/svg&gt;  
&lt;/span&gt;&lt;/template&gt;

|  | &lt;span class=&quot;pl-smi&quot;&gt;county.2nd&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; merge(&lt;span class=&quot;pl-smi&quot;&gt;county.1st&lt;/span&gt;, &lt;span class=&quot;pl-smi&quot;&gt;nc.neighbors&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;by.x&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;firstiterationfirstdegree&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;, &lt;span class=&quot;pl-v&quot;&gt;by.y&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;=&lt;/span&gt; &lt;span class=&quot;pl-s&quot;&gt;&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;county&lt;span class=&quot;pl-pds&quot;&gt;&quot;&lt;/span&gt;&lt;/span&gt;) |
|---|---|
|  |  |
|  | &lt;span class=&quot;pl-c&quot;&gt;&lt;span class=&quot;pl-c&quot;&gt;\#&lt;/span&gt; — find second degree neighbors that are not in first degree neighbor set &lt;/span&gt; |
|  | &lt;span class=&quot;pl-smi&quot;&gt;county.2nd.only&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;county.2nd&lt;/span&gt;\[&lt;span class=&quot;pl-k&quot;&gt;!&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;county.2nd&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;firstdegreeneighbors&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;%in%&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;county.1st&lt;/span&gt;&lt;span class=&quot;pl-k&quot;&gt;$&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;firstiterationfirstdegree&lt;/span&gt;,\] |
|  | &lt;span class=&quot;pl-smi&quot;&gt;county.2nd.only&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;&amp;lt;-&lt;/span&gt; subset(&lt;span class=&quot;pl-smi&quot;&gt;county.2nd.only&lt;/span&gt;, &lt;span class=&quot;pl-k&quot;&gt;!&lt;/span&gt;&lt;span class=&quot;pl-smi&quot;&gt;firstdegreeneighbors&lt;/span&gt; &lt;span class=&quot;pl-k&quot;&gt;%in%&lt;/span&gt; &lt;span class=&quot;pl-smi&quot;&gt;county.of.interest&lt;/span&gt;) |

&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class=&quot;gist-meta&quot;&gt; [view raw](https://gist.github.com/a8dx/6d2e63a87a95b57af61ddb6b19fcf936/raw/d649cb190edea524106324478eb6b93670463a7d/SecondDegreeNeighbors.r)  
 [  
 SecondDegreeNeighbors.r  
 ](https://gist.github.com/a8dx/6d2e63a87a95b57af61ddb6b19fcf936#file-seconddegreeneighbors-r)  
 hosted with ❤ by [GitHub](https://github.com) &lt;/div&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;Now let’s plot. If we use the basic &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plot&lt;/code&gt; command, we can create a relatively decent looking figure without much work, with Chatham County in purple, its first-degree neighbors in red and its second-degree neighbors in blue.&lt;br /&gt;
&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/01/nthdegreeneighborsbase.png?resize=712%2C267&amp;amp;ssl=1&quot; alt=&quot;NthDegreeNeighborsBase&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Alternatively we can convert our &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;sf&lt;/code&gt; objects to the Spatial class, and then layer them onto a Leaflet basemap. This map uses the &lt;a href=&quot;https://rstudio.github.io/leaflet/basemaps.html&quot;&gt;CartoDB positron tiling&lt;/a&gt;, which provides a nice background, especially when visualizing land areas bordering large bodies of water.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/01/chatham_nc_neighbors.jpeg?resize=680%2C264&amp;amp;ssl=1&quot; alt=&quot;Chatham_NC_Neighbors&quot; /&gt;&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>andagostino</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      
        <category term="R" />
      
        <category term="Visualization" />
      

      

      
        <summary type="html">You’re more likely to complain about the neighbors upstairs who are making noise after midnight than those in an apartment two buildings away. Proximity matters and that’s patently obvious, but oftentimes it takes a bit of work to identify who is close and who isn’t. While raster data is packaged in a consistent gridded format for which inverse distance weighting schemes readily can be applied, shapefiles with oddly-shaped features, like these gerrymandered districts, may present more of a challenge. Fortunately the simple features library in R can save the day and with little sweat on your brow.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Local Macros in Stata Using Regular Expressions</title>
      
      
      <link href="https://anthonylouisdagostino.com/local-macros-in-stata-using-regular-expressions/" rel="alternate" type="text/html" title="Local Macros in Stata Using Regular Expressions" />
      
      <published>2017-12-13T19:21:12+00:00</published>
      <updated>2017-12-13T19:21:12+00:00</updated>
      <id>https://anthonylouisdagostino.com/local-macros-in-stata-using-regular-expressions</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/local-macros-in-stata-using-regular-expressions/">&lt;p&gt;Regular expressions can dramatically make your scripting simpler, more automated, and enable you to embed systematically-important information in filenames, variables, dictionaries, and paths. With enough practice, xkcd reminds us that regexp can also &lt;a href=&quot;https://xkcd.com/208/&quot;&gt;make you a superhero&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Stata provides a &lt;a href=&quot;https://www.stata.com/support/faqs/data-management/regular-expressions/&quot;&gt;very nice table&lt;/a&gt; of their regular expressions and offers some helpful examples, but these seem more geared towards creating derivative variables, like extracting the area code from a telephone number string variable. Other objectives require a different tack.&lt;/p&gt;

&lt;p&gt;I often want to use regular expressions to make a local that can be manipulated on the fly. Consider scanning through a vector of variable names and extracting some key feature, like a two-letter state abbreviation or a four digit year. In such a case, I don’t want to create a new variable (it’d hold the same value across all rows anyway), so the standard &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;gen newvar = something&lt;/code&gt; approach isn’t going to apply here.&lt;/p&gt;

&lt;p&gt;When exporting weather data, I’ll save temperature-specific values in a unique variable. For example, the total number of days in a season that a given pixel is exposed to temperatures between 23.5ºC and 24.5ºC might be named &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;tempExpos_24&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;We could then run something like,&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;if regexm(&quot;`x&apos;&quot;, &quot;([0-9][0-9])&quot;) { 
    loc degree = &quot;`=regexs(1)&apos;&quot;
    }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;to extract the temperature, and then do anything from constructing new variables that use that value, renaming the variable (e.g., converting ºC to K), or exporting to file with a meaningful name.&lt;/p&gt;

&lt;p&gt;Another possibility doesn’t directly require an if statement, but follows immediately after a display. Given the &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;fn&lt;/code&gt; local which takes a filename as a string, a regexp local can be made with syntax like,&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;di regexm(&quot;`fn&apos;&quot;, &quot;(_ppt_)([0-9][0-9][0-9][0-9])(_)([a-z0-9]*)(_)&quot;)  
    loc year = regexs(2)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;These are departures from the more usual regexp application of generating new variables from other variables, like in this example,&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;gen dum_val = regexs(3) if regexm(parm, &quot;(_)(tmaxpop_)([0-9]*)(_)&quot;)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;which followed a &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;parmest&lt;/code&gt; dump of model estimates.&lt;/p&gt;

&lt;p&gt;And lastly, a more complete example of how multiple locals can be used in an iterative process with regular expressions. What this code snippet does is scan through a locally-defined range of variables, and constructs an aggregate degree day measure relative to an adjustable lower bound (&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;y&lt;/code&gt;). Since this is a composite function requiring many adjacent variables, it is continuously redefining itself until it reaches an upper value, which are the dataset-specified &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;*_max&lt;/code&gt; variables that precede the snippet.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;// manually-identified data-specific season temperature max 
loc ERA_Kh_max = 44
loc ERA_Rb_max = 36 

loc IMD_Kh_max = 40
loc IMD_Rb_max = 34 

loc APHRO_Kh_max = 42
loc APHRO_Rb_max = 35


** Construct cumulative harmful degree day totals for a range of potential cut-off points 
foreach sea in &quot;Kh&quot; &quot;Rb&quot; {  
    foreach z in  &quot;ERA&quot; &quot;IMD&quot; &quot;APHRO&quot; { 
        forval y = 26(1)31 { 
            gen `z&apos;_`sea&apos;_`y&apos;plus_tmeanDD_Sum = 0 
                lab var `z&apos;_`sea&apos;_`y&apos;plus_tmeanDD_Sum &quot;Cumulative degree days &amp;gt;= `y&apos; for `sea&apos; Season, using `z&apos; Temp Data [tmean DD approach]&quot;
            
            // collect all variables for the specified temperature range 
            unab `z&apos;`y&apos;plus: `z&apos;_tmeandd_`sea&apos;_`y&apos; - `z&apos;_tmeandd_`sea&apos;_``z&apos;_`sea&apos;_max&apos; 
                foreach x in ``z&apos;`y&apos;plus&apos; { 
                    if regexm(&quot;`x&apos;&quot;, &quot;([0-9][0-9])&quot;) { 
                        loc degree = &quot;`=regexs(1)&apos;&quot;
                    }
                    
                    loc degint = int(`degree&apos;)
                    replace `z&apos;_`sea&apos;_`y&apos;plus_tmeanDD_Sum = `z&apos;_`sea&apos;_`y&apos;plus_tmeanDD_Sum + ((`degint&apos; - `y&apos;) *  `x&apos;)               
                } 
        } 
    }   
}   
          
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;</content>

      
      
      
      
      

      
        <author>
            <name>andagostino</name>
          
          
        </author>
      

      
        <category term="climate change" />
      
        <category term="Computing" />
      

      

      
        <summary type="html">Regular expressions can dramatically make your scripting simpler, more automated, and enable you to embed systematically-important information in filenames, variables, dictionaries, and paths. With enough practice, xkcd reminds us that regexp can also make you a superhero.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Converting .TXT/.GRD Climate Data Files to netCDF Format</title>
      
      
      <link href="https://anthonylouisdagostino.com/converting-txt-grd-climate-data-to-netcdf/" rel="alternate" type="text/html" title="Converting .TXT/.GRD Climate Data Files to netCDF Format" />
      
      <published>2017-10-08T03:49:22+00:00</published>
      <updated>2017-10-08T03:49:22+00:00</updated>
      <id>https://anthonylouisdagostino.com/converting-txt-grd-climate-data-to-netcdf</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/converting-txt-grd-climate-data-to-netcdf/">&lt;p&gt;Climate data is packaged and distributed in &lt;del&gt;too&lt;/del&gt; many file formats. Under ideal circumstances, you could easily convert data from formats you’re not familiar with (and don’t have scripts to handle), to those that you do. This is why analogous tools like &lt;a href=&quot;http://www.stattransfer.com/&quot;&gt;Stat/Transfer&lt;/a&gt; for statistical databases often used by social scientists, are so helpful. If a stranger on the street gives you SPSS data, you can on-the-fly convert it to something which is Stata-readable. Albeit, the value of software like Stat/Transfer diminishes as more stat packages have comprehensive in-built conversion tools, like R’s &lt;a href=&quot;https://cran.r-project.org/web/packages/readstata13/README.html&quot; target=&quot;_blank&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;readstata13&lt;/code&gt;&lt;/a&gt; and &lt;a href=&quot;http://pandas.pydata.org/pandas-docs/version/0.20/generated/pandas.read_stata.html&quot; target=&quot;_blank&quot;&gt;&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;read stata&lt;/code&gt;&lt;/a&gt; in pandas. Getting similar functionality with climate data requires a bit more lift.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;This post is a tutorial and link to scripts that can convert the .TXT/.GRD file combination format used by the [India Meteorological Department &lt;a href=&quot;http://www.imd.gov.in/&quot; target=&quot;_blank&quot;&gt;IMD&lt;/a&gt; into formats that are more usable for people working with climate data. And if you’re not trained in the sciences, but rather as an economist, figuring out how to use this data often comes with its own frustrations. I hope this helps.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;I rely on the [Climate Data Operators &lt;a href=&quot;https://code.mpimet.mpg.de/projects/cdo&quot; target=&quot;_blank&quot;&gt;CDO&lt;/a&gt; to do the heavy lifting in my climate data workflow and work almost exclusively with the &lt;a href=&quot;https://www.unidata.ucar.edu/software/netcdf/&quot; target=&quot;_blank&quot;&gt;netCDF4 file format.&lt;/a&gt; Since IMD provides their climate data in .TXT/.GRD format, extra work is required to turn those files into more familiar formats. Here we’ll convert it to netCDF, which after processing can then be exported to a spreadsheet.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;h2 id=&quot;data-setup-and-processing-steps&quot;&gt;Data Setup and Processing Steps&lt;/h2&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;

&lt;p&gt;IMD provides you daily data for each weather variable (TMAX, TMIN, TAVG) in year-specific .TXT and .GRD files. The .TXT file includes the lat/lon grid boundaries, timestamp, and daily values for each pixel. The .GRD file somehow converts this long .TXT into an array stack.&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Step 1 – We first generate year-specific .CTL files which contain the data header, since IMD provides us only a single .CTL. I found doing this in R to be relatively straightforward, since a .CTL can be read as a standard text file. Each of the generated .CTL files designate the source data and the spatial/temporal grid resolution. Here I use a modulo operator to differentiate leap years and accordingly modify the number of time-steps. If leap year date data is not included, then this wouldn’t be a concern. The following R code is an example of how those .CTLs can be auto-generated for a specified year range.&lt;br /&gt;
&lt;br /&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;div class=&quot;language-r highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Filename: convert_to_GrADS.R&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Author: Anthony Louis D&apos;Agostino (ald at stanford dot edu)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Date Created: September 16, 2015&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Last Edited: October 07, 2017&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Purpose: Modifies the NCC-provided CTL file and generates unique versions for each year of data, each type of temperature&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# max, min, mean.&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# Notes: While perhaps an overkill, this is generalizable to a setting where grids are file-specific.&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;rm&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;list&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ls&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;())&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;root.path&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/tmp&quot;&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- create path for generated GrADS control .CTL files.&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl.path&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;root.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;CTL_Files&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sep&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;file.exists&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;){&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dir.create&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- data stored in three separate variable-specific folders&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;temp_types&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;nf&quot;&gt;c&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;MinT&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;MaxT&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;MeanT&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;in&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;temp_types&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;type.path&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sep&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;file.exists&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;type.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;){&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;dir.create&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;type.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- year range for which data is available&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;for&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;in&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;1951&lt;/span&gt;&lt;span class=&quot;o&quot;&gt;:&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;2014&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;){&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;print&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;Now processing year &quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot; for variable &quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;))&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- read in and update template .CTL provided by IMD&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;read.delim&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;root.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/Temp.ctl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;header&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sep&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot; &quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;stringsAsFactors&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- file reference for source GRD&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;1&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;V2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;root.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;_&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;.GRD&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;if&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;%%&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;4&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;==&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;){&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;7&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;V2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;366&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# account for leap years in total number of timesteps&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;k&quot;&gt;else&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;{&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
      &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;7&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;V2&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;m&quot;&gt;365&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
    &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;[&lt;/span&gt;&lt;span class=&quot;m&quot;&gt;7&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;V5&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;]&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;&amp;lt;-&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;1JAN&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;c1&quot;&gt;# -- write updated .CTL to file&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
  &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;write.table&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;ctl_file&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;type.path&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;paste0&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;(&lt;/span&gt;&lt;span class=&quot;n&quot;&gt;x&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;_&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;i&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;.ctl&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;sep&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;/&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;),&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;quote&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;col.names&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;na&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;s2&quot;&gt;&quot;&quot;&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;,&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;n&quot;&gt;row.names&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;o&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;w&quot;&gt; &lt;/span&gt;&lt;span class=&quot;kc&quot;&gt;FALSE&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;)&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

  &lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;
&lt;/span&gt;&lt;span class=&quot;p&quot;&gt;}&lt;/span&gt;&lt;span class=&quot;w&quot;&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;
&lt;p&gt;&lt;br /&gt;&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;Step 2 - CDO enables easy conversion between GrADS and netCDF which we’ll exploit. This could be done at the command line, for reproducibility we’ll use the Python wrappers which can be installed via pip. The following script demonstrates how those binaries can be read in, converted to netCDF, and concatenated by temperature variable.&lt;br /&gt;
&lt;br /&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;```python&lt;/p&gt;

&lt;h1 id=&quot;filename-imd_convertrawdatapy&quot;&gt;Filename: IMD_convertRawData.py&lt;/h1&gt;
&lt;h1 id=&quot;author-anthony-louis-dagostino-ald-at-stanford-dot-edu&quot;&gt;Author: Anthony Louis D’Agostino (ald at stanford dot edu)&lt;/h1&gt;
&lt;h1 id=&quot;date-created-06012017&quot;&gt;Date Created: 06/01/2017&lt;/h1&gt;
&lt;h1 id=&quot;last-edited-10072017&quot;&gt;Last Edited: 10/07/2017&lt;/h1&gt;
&lt;h1 id=&quot;data-from-ncc-zip-file&quot;&gt;Data: from NCC ZIP file&lt;/h1&gt;
&lt;h1 id=&quot;purpose-reads-in-grads-data-files-concatenates-them-and-then-exports-netcdf-versions&quot;&gt;Purpose: Reads in GRaDS data files, concatenates them, and then exports netCDF versions&lt;/h1&gt;
&lt;h1 id=&quot;notes-to-be-run-after-convert_to_gradsr&quot;&gt;Notes: To be run after “convert_to_GrADS.R”&lt;/h1&gt;

&lt;p&gt;import os
from cdo import *
from netCDF4 import Dataset
import numpy as np
#import pandas as pd&lt;/p&gt;

&lt;p&gt;cdo = Cdo()&lt;/p&gt;

&lt;p&gt;root_path = “/tmp”
ctl_root = os.path.join(root_path, “CTL_Files”)&lt;/p&gt;

&lt;p&gt;def tempOutput(var, ctl, root):
	“””
	Read in binary data and output as a netCDF file.
	var: weather variable &lt;br /&gt;
	ctl: root path for all .ctl’s
	root: project root path
	“””&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;# -- rename variables for consistency with other projects
temp_rename = {&quot;MaxT&quot;: &quot;tmax&quot;, &quot;MinT&quot;: &quot;tmin&quot;, &quot;MeanT&quot;: &quot;tmean&quot;}

# initialize using first year&apos;s data
t = cdo.import_binary(input = os.path.join(ctl, var, var + &quot;_1951.ctl&quot;))

# -- loop through each remaining year
for y in range(1952,2015):
	print &quot;Now processing &quot; + str(y)
	fn = var + &quot;_&quot; + str(y) + &quot;.ctl&quot;
	print &quot;Processing &quot; + str(os.path.join(ctl_root, var, fn))

	# -- concatenate into a single file
	data = cdo.import_binary(input = os.path.join(ctl_root, var, fn))
	t = cdo.cat(input = &quot; &quot;.join([t,data]))

# -- save variable-specific file
t = cdo.copy(input = t, options = &quot;-f nc&quot;, output = os.path.join(root, temp_rename[var] + &quot;Proc.nc&quot;))
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;tempOutput(“MaxT”, ctl_root, root_path)
tempOutput(“MinT”, ctl_root, root_path)
tempOutput(“MeanT”, ctl_root, root_path)&lt;/p&gt;

&lt;p&gt;```&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Step 3 – Now the concatenated output file has been saved as a netCDF, which means you can perform all the standard CDO operators on it. You can also simply read your files in R, and run functions like spatial averaging with a minimum of code, as in &lt;a href=&quot;https://gis.stackexchange.com/questions/213493/area-weighted-average-raster-values-within-each-spatialpolygonsdataframe-polygon&quot; target=&quot;_blank&quot;&gt;this example&lt;/a&gt;.&lt;br /&gt;
 &lt;br /&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As always, happy to field any questions you might have on the code and workflow!&lt;br /&gt;
&lt;br /&gt;&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>Anthony Louis D&apos;Agostino</name>
          
          
        </author>
      

      
        <category term="climate change" />
      
        <category term="computing" />
      
        <category term="data" />
      
        <category term="India" />
      
        <category term="Python" />
      
        <category term="R" />
      

      

      
        <summary type="html">Climate data is packaged and distributed in too many file formats. Under ideal circumstances, you could easily convert data from formats you’re not familiar with (and don’t have scripts to handle), to those that you do. This is why analogous tools like Stat/Transfer for statistical databases often used by social scientists, are so helpful. If a stranger on the street gives you SPSS data, you can on-the-fly convert it to something which is Stata-readable. Albeit, the value of software like Stat/Transfer diminishes as more stat packages have comprehensive in-built conversion tools, like R’s readstata13 and read stata in pandas. Getting similar functionality with climate data requires a bit more lift.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Visualizations Gone Wild</title>
      
      
      <link href="https://anthonylouisdagostino.com/visualizations-gone-wild/" rel="alternate" type="text/html" title="Visualizations Gone Wild" />
      
      <published>2017-10-06T01:13:00+00:00</published>
      <updated>2017-10-06T01:13:00+00:00</updated>
      <id>https://anthonylouisdagostino.com/visualizations-gone-wild</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/visualizations-gone-wild/">&lt;blockquote&gt;
  &lt;p&gt;Who said art had to be intentional?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I don’t spend as much time making art as I would like, so it’s always a pleasant surprise when bad code inadvertently helps me fix that.&lt;/p&gt;

&lt;h2 id=&quot;under-a-yellowcake-sun&quot;&gt;Under a Yellowcake Sun&lt;/h2&gt;

&lt;p&gt;This is “Desert Mountains, Sunken Sun” and is brought to you by &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;ggplot2&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2017/10/desertmountainsandsun.png?resize=487%2C421&amp;amp;ssl=1&quot; alt=&quot;DesertMountainsandSun&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Using&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stat_binhex()&lt;/code&gt;you can create colorful heat map density plots with hexagons representing a collection of observations with adjacent values. There is a problem with stretch, such that hexagons will becomes exceptionally large if few observations with comparable values exist.&lt;/p&gt;

&lt;p&gt;What’s the quick solution? Toggle the number of bins over which the data is represented using &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stat_binhex(bins = n)&lt;/code&gt;, where the default &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;n&lt;/code&gt; is 30, and consider applying a log or log10 transform (e.g., &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;trans = &quot;log10&quot;&lt;/code&gt; in your &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;scale_fill_XXXX&lt;/code&gt; command) which controls the values that correspond to the legend’s tick marks. You’ll likely have to toy around with a few combinations before finding the most visually informative values.&lt;/p&gt;

&lt;h2 id=&quot;a-fragmented-country&quot;&gt;A Fragmented Country&lt;/h2&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2018/07/pm2p5_zcta5_anybreaks-e1531162151710.png?resize=712%2C448&amp;amp;ssl=1&quot; alt=&quot;PM2p5_ZCTA5_AnyBreaks&quot; /&gt;&lt;/p&gt;

&lt;p&gt;This was a byproduct of a poorly executed filtering procedure using a &lt;a href=&quot;https://www.census.gov/geo/maps-data/data/cbf/cbf_zcta.html&quot;&gt;Census ZCTA5 shapefile&lt;/a&gt;. &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;plot(st_geometry(object[&quot;column&quot;))&lt;/code&gt; helped drive us home.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      
        <category term="R" />
      

      

      
        <summary type="html">Who said art had to be intentional?</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Recent Trends in Women’s Employment in Rural India</title>
      
      
      <link href="https://anthonylouisdagostino.com/recent-trends-in-womens-employment-in-rural-india/" rel="alternate" type="text/html" title="Recent Trends in Women&apos;s Employment in Rural India" />
      
      <published>2017-10-01T06:29:22+00:00</published>
      <updated>2017-10-01T06:29:22+00:00</updated>
      <id>https://anthonylouisdagostino.com/recent-trends-in-womens-employment-in-rural-india</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/recent-trends-in-womens-employment-in-rural-india/">&lt;p&gt;That female labor force participation (FLFP) is U-shaped in per capita income is one significant stylized fact at the intersection of development and labor economics. At low levels of development, subsistence requirements render women’s work a necessity for household survival. At higher incomes, the nature of employment changes, with the growth of manufacturing jobs that tend to be staffed by men, and may accompany a contraction of agricultural labor. This, as well as social norms like purdah which encourage women’s non-participation in paid work or their association with non-household men, squeeze FLFP. Further along the industrialization growth path, these manufacturing jobs pave the way for service sector employment, for which women may exercise a comparative advantage. High wage jobs also induce higher FLFP through an opportunity cost channel. Staying home could mean a lot of foregone household income.&lt;/p&gt;

&lt;p&gt;By no means is this law-like. Some recent empirical work, including by &lt;a href=&quot;https://link.springer.com/article/10.1007/s00148-013-0488-2&quot;&gt;Gaddis and Klasen&lt;/a&gt;, calls into question the U-shape’s robustness, and while &lt;a href=&quot;http://www.nber.org/papers/w22766&quot;&gt;Heath and Jayachandran&lt;/a&gt; document the continued presence of the general shape, they also find the curve shifting upwards. This suggests that for a given level of real per capita income, women’s employment rates on a global basis have risen over time.&lt;/p&gt;

&lt;p&gt;In India there’s been growing concern from various quarters that advancement in women’s labor force participation is not only stalling, but &lt;a href=&quot;http://www.bbc.com/news/world-asia-india-39945473&quot;&gt;reversing&lt;/a&gt;. Since per capita income levels lie to the left of the U’s inflection point, then we should anticipate further declines. One peculiarity between the stylized observation and the results from India is the discrepancy in levels. For example, Heath and Jayachandran estimate a minimum FLFP of around 40%, yet we observe levels even lower than this in rural India for some rounds of the National Sample Survey’s Employment &amp;amp; Unemployment surveys, likely the best source of employment data available in India.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2017/11/nss_flfpunweighted1.png?resize=712%2C545&amp;amp;ssl=1&quot; alt=&quot;NSS_FLFPUnWeighted&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Digging deeper one layer into these employment trends reveals that employment shares by wage type have been relatively consistent across rounds, with the exception of 1987. This may have something to do with Round 43’s survey design question on primary status applying to the preceding 7 day period, whereas latter rounds used a 365 day period. In each survey year, about 40% of rural women were engaged in paid labor, with the remainder either working on an unpaid basis or ‘self-employed’ in a family business or farm operation.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2017/11/ruralemptype2.png?resize=712%2C545&amp;amp;ssl=1&quot; alt=&quot;RuralEmpType&quot; /&gt;&lt;/p&gt;

&lt;p&gt;We can also examine the &lt;a href=&quot;http://mail.mospi.gov.in/index.php/catalog/143/study-description&quot;&gt;NSS 68th Round&lt;/a&gt; of Employment &amp;amp; Unemployment data for 2011/2012, which is the most recent round for which this module was conducted. As before, I restrict the sample to rural women aged 15-65, leaving a sample size of nearly 93,000. As shown in Table 1 below, 3 out of 5 women are primarily engaged in unpaid domestic work, or that in conjunction with free collection of items like firewood and water, or activity like sewing and tailoring for household members. For ease, let’s refer to the combination of these two activities as Housework+.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2017/11/prinstatus.png?resize=712%2C535&amp;amp;ssl=1&quot; alt=&quot;PrinStatus&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The principal (primary) activity paints an incomplete picture, especially when women work part-time or seasonally and therefore likely to respond as principally involved in domestic work. An alternative is to also consider their subsidiary activity, and take the minimum value of the two. If their secondary status involves work of a non-domestic nature, a respondent is effectively pushed into an upper category. This approach works because NSS formats the values such that active labor force participation responses precede those of unpaid domestic work. As an example, a woman who reports as primarily a student and secondarily as a casual wage laborer, would be coded as the latter in Table 2.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://i0.wp.com/anthonylouisdagostino.com///wp-content/uploads/2017/11/subsidstatus.png?resize=712%2C441&amp;amp;ssl=1&quot; alt=&quot;SubsidStatus&quot; /&gt;&lt;/p&gt;

&lt;p&gt;The share of women involved in the Housework+ categories drop to below 50%, with Self-Employed and Unpaid Family Workers the largest gaining categories. While these women may primarily be engaged in domestic work, a portion of their time is also tied to the family business or farm, with no guarantee of direct wages from their labor. All the ‘working’ categories still sum up to only 35% of the population, though is a sizable improvement on the 25% when counting only reported principal activities.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;&lt;em&gt;Data used for the 1983-2009 analysis comes from &lt;a href=&quot;https://doi.org/10.18128/D020.V6.5&quot;&gt;IPUMS&lt;/a&gt;, which sources the NSS Employment survey data from the Ministry of Statistics and Programme Implementation, India (MOSPI). Data from the 2011 analysis comes directly from MOSPI.&lt;/em&gt;&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Gender" />
      

      

      
        <summary type="html">That female labor force participation (FLFP) is U-shaped in per capita income is one significant stylized fact at the intersection of development and labor economics. At low levels of development, subsistence requirements render women’s work a necessity for household survival. At higher incomes, the nature of employment changes, with the growth of manufacturing jobs that tend to be staffed by men, and may accompany a contraction of agricultural labor. This, as well as social norms like purdah which encourage women’s non-participation in paid work or their association with non-household men, squeeze FLFP. Further along the industrialization growth path, these manufacturing jobs pave the way for service sector employment, for which women may exercise a comparative advantage. High wage jobs also induce higher FLFP through an opportunity cost channel. Staying home could mean a lot of foregone household income.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">San Francisco-Stanford Commuting</title>
      
      
      <link href="https://anthonylouisdagostino.com/san-francisco-stanford-commuting/" rel="alternate" type="text/html" title="San Francisco-Stanford Commuting" />
      
      <published>2017-08-24T03:55:48+00:00</published>
      <updated>2017-08-24T03:55:48+00:00</updated>
      <id>https://anthonylouisdagostino.com/san-francisco-stanford-commuting</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/san-francisco-stanford-commuting/">&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2017/08/rear_de.jpg?w=199&amp;amp;resize=199%2C300&quot; alt=&quot;rear_de&quot; /&gt;I’ve now been commuting from San Francisco to Stanford for nearly four weeks and thought it might be helpful to pen some observations for folks who are considering living in SF, but working in Palo Alto or at Stanford (at each of the postdoc activities I’ve attended, we’ve been reminded there are more than 2,300 of us campus-wide at a time, so I think there’s an audience for this). Prior to moving here, I tried to research the viability of the commute and how it would work with biking, but found the level of detail lacking on posts at Quora and elsewhere. In sum, it’s entirely doable with bikes and haven’t gotten too exhausted by it yet.&lt;/p&gt;

&lt;p&gt;My preferred commute consists of a couple separate legs. First, I have an SF-based bike that I ride from my apartment in SoMa to the 4th and King Street Caltrain station. This is about 6 minutes, and there’s a bike path the entire length of Townsend Street. I then leave this bike at the &lt;a href=&quot;http://bikehub.com/caltrain-bike-station/&quot;&gt;staffed Bike Station parking facility&lt;/a&gt;, which remains open until 8:30PM. Parking there is free through the first night, but subsequent nights are charged at $5 each. The Bike Station also can repair your bike while you’re at work, and they have a small shop with some basic accessories like fenders, tubes, and tires for sale. An alternative to parking here would be renting out a Caltrain bike locker, but currently &lt;a href=&quot;http://www.caltrain.com/riderinfo/Bicycles/BicycleParking.html&quot;&gt;all 180 lockers are rented out.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the Palo Alto side, I keep a bike parked overnight at the &lt;a href=&quot;http://home.bikestation.com/bikestation-palo-alto&quot;&gt;Bikestation&lt;/a&gt; which operates a large room accessible by keyless fob on the Southbound side of the train station. There’s a sufficient number of racks and security so far seems to be quite good (i.e., make sure no one follows you into the room without using their fob). You can get 12-months of parking access for slightly more than $100, though some people have suffered bad experiences with inadequate reminder notices informing them of impending contract expiration, and &lt;a href=&quot;https://www.yelp.com/biz/bikestation-palo-alto-palo-alto&quot;&gt;have been locked out for several days on end despite renewing.&lt;/a&gt; This leg, from Palo Alto station to Stanford office bike rack, takes about 8 minutes and involves only 3 major intersections where you could potentially be delayed.&lt;/p&gt;

&lt;p&gt;For several weeks I did the standard haul-your-bike-on-the-train routine, park it in one of two bike cars (if you’re riding the Baby Bullet), and hopefully find a nearby seat. I never got bumped from a train because of too few rack spaces available, but you will spend a decent amount of time organizing your bike, moving it around, loading and unloading off the train, and waiting for others. Since there’s a sizable crowd traveling from 4th/King to Palo Alto, you can usually find a stack of bikes heading to your destination, so there’s no need to reorganize to ensure the nearest bike belongs to the first person disembarking.&lt;/p&gt;

&lt;p&gt;A major downside to this commute is that a lot of folks are relatively clueless or indifferent about how they treat other peoples’ bikes (and even their own), so doing this day in day out is likely to cause some spokes to get pulled and for your derailleurs to get caught up in another bike’s components. Better to bring a junker, even at the risk of your cyclist cred being compromised. Your chances of finding both a seat and a spot to place your bike are obviously increasing in trains traveling at off-peak hours, but given that conductors usually turn a blind eye to 5 and maybe even 6 bikes on a stack (when 4 is Caltrain’s official maximum), you needn’t worry about finding a space on the morning southbound if you’re boarding at 4th/King. There is a second wave of cyclists on-boarding at 22nd in Dogpatch, and they inevitably have a tougher time finding space and organizing their bikes, especially those who are heading further south than Mountain View. So far, I haven’t seen those folks turned away either, but it may be happening on earlier trains than mine.&lt;/p&gt;

&lt;p&gt;In total, the door-to-door trip is slightly less than an hour and ten minutes, in large part from living within a few blocks of the SF station. I’d say the absolute commute minimum would be living at Avalon, skipping the SF bike portion, and simply walking to the morning train. Alternatively, Potrero Hill may be an option for you because of its proximity to the 22nd Street station, but the hilliness is hit-or-miss and what appears to be close could be quite a time-consuming walk or ride. Overall, the availability of amenities and density of grocery stores, restaurants, retail is also substantially higher in SoMa, providing more reasons to opt for this part of town where there’s simply more going on, at both higher cost and more limited green spaces.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Transportation" />
      

      

      
        <summary type="html">I’ve now been commuting from San Francisco to Stanford for nearly four weeks and thought it might be helpful to pen some observations for folks who are considering living in SF, but working in Palo Alto or at Stanford (at each of the postdoc activities I’ve attended, we’ve been reminded there are more than 2,300 of us campus-wide at a time, so I think there’s an audience for this). Prior to moving here, I tried to research the viability of the commute and how it would work with biking, but found the level of detail lacking on posts at Quora and elsewhere. In sum, it’s entirely doable with bikes and haven’t gotten too exhausted by it yet.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Effortlessly Merging 1,000s of Raw Data Files with Stata</title>
      
      
      <link href="https://anthonylouisdagostino.com/effortlessly-merging-1000s-of-raw-data-files-with-stata/" rel="alternate" type="text/html" title="Effortlessly Merging 1,000s of Raw Data Files with Stata" />
      
      <published>2017-03-30T03:46:16+00:00</published>
      <updated>2017-03-30T03:46:16+00:00</updated>
      <id>https://anthonylouisdagostino.com/effortlessly-merging-1000s-of-raw-data-files-with-stata</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/effortlessly-merging-1000s-of-raw-data-files-with-stata/">&lt;p&gt;I frequently have to consolidate 100s or 1,000s of raw data files into Stata, that are potentially stored in numerous and potentially unknown folders and subfolders, and have developed a workflow that I think is useful. This approach means I don’t need to know file names or paths, and can instead assign search parameters that determine which files get tagged for processing. I’ve used this approach when constructing a financial flows panel database for India from numerous state-level bank deposits spreadsheets, as well as manipulating GCM output that was chopped up and spit out into thousands of files using an Python/ArcPy workflow. I frequently encounter these setups, and if you do too then this tutorial is probably relevant. All the files needed to work through the following example are on &lt;a href=&quot;https://github.com/a8dx/Stata-Raw-Data-Merge&quot;&gt;github&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;philosophy&quot;&gt;Philosophy&lt;/h2&gt;

&lt;p&gt;This merge procedure satisfies two key requirements:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;Agile – hard-coding is either minimized or non-existent, which means you can update folder contents, including the addition of new subfolders, and the script will automatically adjust. Agility would not be achieved, for example, by using “medium-coding” solutions (somewhere in the middle of soft and hard) like for loops with fixed end points, since changes to the range do not necessarily alter the for loop structure. To be concrete, assume someone from the House Intelligence Committee handed you files that spanned “WiretappingRawData_32.csv” to “WiretappingRawData_56.csv.” You could write a for loop spanning 32 to 56, but this does nothing when files *_57 onwards are added.&lt;/li&gt;
  &lt;li&gt;Versatile – can easily be adapted to a range of data formats and processing requirements. The major code chunks below can easily accommodate CSVs, XLSXs, binary files, etc., and slots into longer Stata workflows.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2 id=&quot;data-context&quot;&gt;Data Context&lt;/h2&gt;

&lt;p&gt;Let’s start with this example, which includes some simplifying elements that we can revisit later. You need to import daily temperature data for several climate models for multiple climate experiments (e.g., &lt;a href=&quot;https://en.wikipedia.org/wiki/Representative_Concentration_Pathways&quot;&gt;representative concentration pathways&lt;/a&gt;) for a single year, which quickly puts us in the realm of 10,000s of files. The directory tree may look like this:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2017/03/screen-shot-2017-03-29-at-9-56-43-pm.png?resize=712%2C996&quot; alt=&quot;Screen Shot 2017-03-29 at 9.56.43 PM&quot; /&gt; where RCP45 and RCP85 respectively refer to output from the RCP 4.5 and 8.5 experiments. All of the 20+ subfolders (ACCESS, BNU, MIROC, etc.) have names that are unknowable in advance, and the lack of a predictable naming scheme necessitates an ‘agile’ approach which organically discovers all the paths. In each subfolder are numerous model-specific files, each containing data for an individual date, like these:&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2017/03/screen-shot-2017-03-29-at-10-10-09-pm.png?resize=712%2C180&quot; alt=&quot;Screen Shot 2017-03-29 at 10.10.09 PM&quot; /&gt;You first need to install &lt;strong&gt;ashell&lt;/strong&gt; which allows us to handily capture and manipulate shell output.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2017/03/screen-shot-2017-03-29-at-10-22-49-pm.png?resize=485%2C67&quot; alt=&quot;Screen Shot 2017-03-29 at 10.22.49 PM&quot; /&gt;&lt;/p&gt;

&lt;p&gt;We’ll use &lt;strong&gt;find&lt;/strong&gt; which is a basic shell command that can be accessed by Stata.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;ashell find &quot;${cmip5Base}&quot; -maxdepth 1 -mindepth 1 -type d
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;This command uses my global ${cmip5Base} folder as the root directory, searches 1 level deep (no more, no less, as dictated by the maxdepth and mindepth parameters) and only searches for paths, denoted by the “-type d” suffix referring to directories, not files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ashell&lt;/strong&gt; will print out the series of paths or files that satisfy the &lt;strong&gt;find&lt;/strong&gt; search, and uses this very helpful indexing system which can be managed with a for loop, where each search return is numerically identified as 1 through `r(no)’, which is the last file. Let’s assume we just want to list the directory contents, and so running the following loop will print out all those paths.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; forvalues j = 1/`r(no)&apos; { 
    loc currentDir &quot;`r(o`j&apos;)&apos;&quot;
    di &quot;Current Dir = `currentDir&apos;&quot; 
 }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Now we want to access the files in each of the &lt;em&gt;j&lt;/em&gt; folders. They are accessible as r-class macros, starting with `r(o1)’ and ending at some J = no. We issue a similar &lt;strong&gt;find&lt;/strong&gt; command, but now stipulate some file parameters.&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; ashell find &quot;`currentDir&apos;&quot; -maxdepth 1 -name &quot;*.csv&quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Instead of the -type option, we’re now using -name with a wildcard to capture only comma-separated files. Keeping maxdepth to 1 restricts the search to only those files stored in `currentDir.’&lt;/p&gt;

&lt;p&gt;This generates another r-class counter spanning ``r(o1)’ to&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt; &lt;/code&gt;r(no)’, which can be handled via a new and differently indexed for loop across &lt;em&gt;i&lt;/em&gt;. Consider this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; forval i = 1/`r(no)&apos; { 
    di &quot;r(o`i&apos;)&quot;
    insheet using &quot;`r(o`i&apos;)&apos;&quot;, comma clear  
 }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Since our target raw data files are CSVs, we need to use the &lt;strong&gt;insheet&lt;/strong&gt; command instead of &lt;strong&gt;import excel.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I then rely on &lt;a href=&quot;http://www.stata.com/support/faqs/data-management/regular-expressions/&quot;&gt;Stata’s extremely helpful regular expressions operators&lt;/a&gt; to extract important details from the filename. In order to do this, you would know ex ante the file naming convention and how to access each segment. As an example, consider a file named&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;tasmax_day_BCSD_rcp85_r1i1p1_CanESM2_2050-01-01_1991District__Power_3.csv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;which tells me which variable (tasmax_day), which experiment (rcpXX), which model (CanESM2), which date (2050-01-01), and which polynomial degree (Power_3) this data is taken from. Ideally, all these pieces would be included in the data itself, but let’s assume they’re not. One way of using regular expressions to strip out those identifiers look something like this, though obviously there are numerous ways you can tackle this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt; * Recover basics from filename, presumably more accurate than if sourced from directory name * 
 gen polynomialOrder = regexs(2) if regexm(&quot;`fn&apos;&quot;, &quot;(Power_)([0-9])(.csv)&quot;)
 gen modelFamily = regexs(2) if regexm(&quot;`fn&apos;&quot;, &quot;(r1i1p1_)([a-zA-Z0-9-]*)(_2050)&quot;)
 gen experiment = regexs(2) if regexm(&quot;`fn&apos;&quot;, &quot;(BCSD_)([a-zA-Z0-9-]*)(_r1)&quot;) 
 gen year = regexs(1) if regexm(dateFixed, &quot;([0-9][0-9][0-9][0-9])(-)&quot;)
 gen month = regexs(2) if regexm(dateFixed, &quot;(-)([0-9][0-9])(-)&quot;)
 gen day = regexs(4) if regexm(dateFixed, &quot;(-)([0-9][0-9])(-)([0-9][0-9])&quot;)
 
 destring, replace
 
 gen date = mdy(month, day, year) 
 format date %td 
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;Let’s say those variables are sufficient to uniquely identify our observations, since we want to merge this particular file with 1000s of others and not worry about those gruesome _merge == 5 errors. This particular spreadsheet is then saved as a temporary file, which can then be seamlessly merged with all others. As an example, we can do this:&lt;/p&gt;

&lt;div class=&quot;language-plaintext highlighter-rouge&quot;&gt;&lt;div class=&quot;highlight&quot;&gt;&lt;pre class=&quot;highlight&quot;&gt;&lt;code&gt;    tempfile temp`i&apos;
    save `temp`i&apos;&apos; 
    di &quot;Just finished saving file `i&apos;&quot; 
   ashell find &quot;`currentDir&apos;&quot; -maxdepth 1 -name &quot;*.csv&quot; 
 } 


use `temp1&apos;, clear 
 
 forval j = 2/`r(no)&apos; { 
        merge 1:1 dist_code date modelFamily experiment using `temp`j&apos;&apos;, update replace
        tab _merge
        drop _merge 
 }
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;

&lt;p&gt;The reason for issuing a second instance of &lt;strong&gt;ashell&lt;/strong&gt; before entering the next iteration of the &lt;em&gt;i&lt;/em&gt; for loop, is because of a quirk which leads the r(o1)…r(no) local index values to be erased when some data processing command is performed. Once the merge is complete, you can then save your master file as a *.dta.&lt;/p&gt;

&lt;p&gt;While the Stata .do I’m sharing explicitly identifies the parent folder, you can change the -maxdepth parameters to search across multiple sets of children paths, with the ability to save all sub-parent folder contents as a separate *.dta. You can then repeat the procedure abov, across your individual *.dta’s, to merge up into a master *.dta. The reason you might wish to adopt an iterative approach is that each subroutine requires lengthy processing time, and you can avoid duplicating work by saving results separately.&lt;/p&gt;

&lt;p&gt;I’d love to hear from you if you found this helpful, or have suggestions for tailoring this to specific workflow requirements you frequently find yourself having to plough through.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      
        <category term="Stata" />
      

      

      
        <summary type="html">I frequently have to consolidate 100s or 1,000s of raw data files into Stata, that are potentially stored in numerous and potentially unknown folders and subfolders, and have developed a workflow that I think is useful. This approach means I don’t need to know file names or paths, and can instead assign search parameters that determine which files get tagged for processing. I’ve used this approach when constructing a financial flows panel database for India from numerous state-level bank deposits spreadsheets, as well as manipulating GCM output that was chopped up and spit out into thousands of files using an Python/ArcPy workflow. I frequently encounter these setups, and if you do too then this tutorial is probably relevant. All the files needed to work through the following example are on github.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Towards Closing Gender Data Gaps</title>
      
      
      <link href="https://anthonylouisdagostino.com/towards-closing-gender-data-gaps/" rel="alternate" type="text/html" title="Towards Closing Gender Data Gaps" />
      
      <published>2016-10-12T01:30:59+00:00</published>
      <updated>2016-10-12T01:30:59+00:00</updated>
      <id>https://anthonylouisdagostino.com/towards-closing-gender-data-gaps</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/towards-closing-gender-data-gaps/">&lt;p&gt;In May, the Bill &amp;amp; Melinda Gates Foundation &lt;a href=&quot;http://www.gatesfoundation.org/Media-Center/Press-Releases/2016/05/Gates-Foundation-Announces-80-Mill-Doll-Comm-Closing-Gender-Data-Gaps-Acc-Progress-for-Women-Girls&quot;&gt;announced a three-year, $80 million investment&lt;/a&gt; towards closing the gender data gap, but I only today came across this &lt;a href=&quot;https://www.youtube.com/watch?v=ekW2U4JoN84&quot;&gt;great video&lt;/a&gt; on the same initiative. A portion of the funds will be directed towards improved data collection, particularly of the time use patterns of women and girls and on household asset ownership inventories. Better data means better information for policymakers conducting program evaluations (i.e., how did that recent cash transfer program differentially impact women’s and men’s employment levels?) and will enable researchers greater insight into the long-run implications of unpaid work (i.e., how does working for a household business affect children’s final education achievement?).&lt;/p&gt;

&lt;p&gt;This will be hugely beneficial and I’m excited about the prospect of more data like this becoming available. My job market paper focuses on the impact of technological change on gender wage inequality in Indian agricultural labor markets, and in the context of the Green Revolution (starting in 1966, and to some extent still in process) I find that one effect of the Green Revolution’s productivity gains was a reduction in women’s participation in agricultural wage labor. There are potentially several reasons why this came about, but consider one strand of ethnographic research which posits that agricultural intensification expands the non-agricultural demands on women’s time, leaving less time for wage labor work. Unfortunately, there isn’t granular time use data during this period to tease out what impacted women are doing and how such changes might affect the household division of labor.&lt;/p&gt;

&lt;p&gt;In another project with several collaborators, I use Indian time use data collected from a pilot study carried out from 1998-1999. While a country-wide survey has recently been collected, the data is not yet available which means that the best understanding anyone has about how Indians, rural and urban, spend their time is nearly 20 years old. A lot has changed since then – the country has become more urban, richer, and educated, with cell phones and computers also playing a large role in what people do, and when they do it. Making such data collection exercises more systematic and routine will offer a very helpful window not only into macro-level questions like how welfare programs affect labor decisions, but also on questions related to exercise, child-care, commuting, and leisure. This investment is welcomed news.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Data" />
      
        <category term="Gender" />
      
        <category term="Policy" />
      

      

      
        <summary type="html">In May, the Bill &amp;amp; Melinda Gates Foundation announced a three-year, $80 million investment towards closing the gender data gap, but I only today came across this great video on the same initiative. A portion of the funds will be directed towards improved data collection, particularly of the time use patterns of women and girls and on household asset ownership inventories. Better data means better information for policymakers conducting program evaluations (i.e., how did that recent cash transfer program differentially impact women’s and men’s employment levels?) and will enable researchers greater insight into the long-run implications of unpaid work (i.e., how does working for a household business affect children’s final education achievement?).</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Stata-Latex esttab Regression Table Output Streamlining</title>
      
      
      <link href="https://anthonylouisdagostino.com/stata-latex-esttab-regression-table-output-streamlining/" rel="alternate" type="text/html" title="Stata-Latex esttab Regression Table Output Streamlining" />
      
      <published>2016-05-27T21:36:45+00:00</published>
      <updated>2016-05-27T21:36:45+00:00</updated>
      <id>https://anthonylouisdagostino.com/stata-latex-esttab-regression-table-output-streamlining</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/stata-latex-esttab-regression-table-output-streamlining/">&lt;p&gt;Researchers spend an excessive amount of time getting up to speed with a field’s chosen tools and methods, excessive because there is often a consensus on best practice and yet those best practices are not made common knowledge. I think the CS and statistics communities have this right in their pushing for open data, transparency, and reproducibility in a way that economics, for example, has been late to the game on. As a result, early-stage PhD students can emulate and save those wasted hours tinkering with multicolumns in Latex or some user unfriendly Stata syntax. I have personally benefited greatly from the likes of &lt;a href=&quot;http://www.jwe.cc/2012/03/stata-latex-tables-estout/&quot;&gt;Jorg Weber&lt;/a&gt; and &lt;a href=&quot;http://www.ats.ucla.edu/stat/stata/faq/estout.htm&quot;&gt;UCLA IDRE&lt;/a&gt;, among the numerous Stack Overflow posts on publishing regression output, and have finally developed a satisfycing Stata-Latex &lt;a href=&quot;http://repec.org/bocode/e/estout/esttab.html&quot;&gt;esttab&lt;/a&gt; workflow which doesn’t require an excessive amount of post-processing in order to be usable. &lt;a href=&quot;http://www.eyalfrank.com/&quot;&gt;Eyal Frank&lt;/a&gt; deserves a hat-tip for helping inspire this process.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://github.com/a8dx/Stata-Tools/blob/master/esttab_Latex_Sample.do&quot;&gt;First check out the sample .do&lt;/a&gt; to follow along, downloadable from Github. It’s a basic Stata function which enables me to change dependent variables (`dv’), but otherwise retain the same specifications across all columns. You’ll notice a few things:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Fixed effects are invoked using estadd and local macros. I follow them with replace to ensure that values correspond to the model that’s just been run, and not a previous model by accident.&lt;/li&gt;
  &lt;li&gt;The row order of fixed effects is specified in estout’s &lt;span style=&quot;text-decoration:underline;&quot;&gt;s&lt;/span&gt;tats parameters (“s(…)”) and need not adhere to the sequence in which they appear following a model.&lt;/li&gt;
  &lt;li&gt;Some manual work is needed to ensure the correct number of columns is used if you want additional header rows, likely when multiple columns have a common dependent variable. In this instance, Columns 1-3 have unconditional days worked as the DV, while Columns 4-6 are conditioned on non-zero outcomes. I’m sure there’s a way to relate {*M} to those values, but I’m fine with manually counting and correcting if you generate a multicolumn width error.&lt;/li&gt;
  &lt;li&gt;If you don’t want this additional header, you can remove the “&amp;amp;multicolumn{3}…” line without problem.&lt;/li&gt;
  &lt;li&gt;I really like \cmidrule, but to prevent those underlines from bleeding into each other, include the (l) argument as seen in \cmidrule(l){2-4}. Tip: \cmidrule can also be a real pain – if you want to use it on a single column, you’ll have to repeat the column value as the second argument, e.g. \cmidrule(l){2-2}, NOT \cmidrule(l){2}. I wasted too much time figuring that out. In order to get midrule to work, load both the booktabs and multirow packages, e.g., \usepackage{booktabs} \usepackage{multirow}.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I really like about the prehead posthead syntax affixed to the bottom of the esttab command is the seamless multicolumn centering of footnote text, which otherwise is likely to misalign at least one column of regression output. All that nonsense is avoided here.&lt;/p&gt;

&lt;p&gt;Of course there’s additional flexibility to add horizontal lines, remove them, insert line breaks, etc., but I think this approach provides a streamlined solution to the problem of efficiently exporting Stata results and still retaining some control over header specifics.&lt;/p&gt;

&lt;p&gt;Lastly, this isn’t possible without esttab, do definitely spend some time investigating the numerous formatting options and flexibility built into the various est* programs.&lt;/p&gt;

&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2016/05/screen-shot-2016-05-27-at-6-13-42-pm.png?resize=712%2C571&quot; alt=&quot;Screen Shot 2016-05-27 at 6.13.42 PM&quot; /&gt;&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      
        <category term="econometrics" />
      
        <category term="LaTeX" />
      

      

      
        <summary type="html">Researchers spend an excessive amount of time getting up to speed with a field’s chosen tools and methods, excessive because there is often a consensus on best practice and yet those best practices are not made common knowledge. I think the CS and statistics communities have this right in their pushing for open data, transparency, and reproducibility in a way that economics, for example, has been late to the game on. As a result, early-stage PhD students can emulate and save those wasted hours tinkering with multicolumns in Latex or some user unfriendly Stata syntax. I have personally benefited greatly from the likes of Jorg Weber and UCLA IDRE, among the numerous Stack Overflow posts on publishing regression output, and have finally developed a satisfycing Stata-Latex esttab workflow which doesn’t require an excessive amount of post-processing in order to be usable. Eyal Frank deserves a hat-tip for helping inspire this process.</summary>
      

      
      
    </entry>
  
  
  
    <entry>
      
      <title type="html">Stata: Union of Macros</title>
      
      
      <link href="https://anthonylouisdagostino.com/stata-union-of-macros/" rel="alternate" type="text/html" title="Stata: Union of Macros" />
      
      <published>2016-03-17T21:40:30+00:00</published>
      <updated>2016-03-17T21:40:30+00:00</updated>
      <id>https://anthonylouisdagostino.com/stata-union-of-macros</id>
      <content type="html" xml:base="https://anthonylouisdagostino.com/stata-union-of-macros/">&lt;p&gt;&lt;img src=&quot;https://pathindependence.files.wordpress.com/2016/03/figure-8-the-ascii-text-stream-produced-when-the-binary-stream-in-is-decompressed-png.jpg?resize=508%2C196&quot; alt=&quot;Figure-8-The-ASCII-text-stream-produced-when-the-binary-stream-in-is-decompressed.png&quot; /&gt;&lt;/p&gt;

&lt;p&gt;Wide data is ok, but I prefer long data any day of the week (at least most days of the week). Data creators/providers may disagree, and in such cases you may have to be creative about how you reshape the data.&lt;/p&gt;

&lt;p&gt;Consider this common scenario: variable suffixes denote some index (e.g., time), but not every variable exists for each index value. For example, you can imagine Stata’s &lt;a href=&quot;http://www.ats.ucla.edu/stat/stata/modules/usesave.htm&quot;&gt;auto&lt;/a&gt; dataset formatted wide with fuel economy, weight, and price each indexed by years spanning 1990-2010 [mpg90 weight90 price90 mpg91 weight91 price91 … mpg10 weight10 mpg10]. To &lt;strong&gt;reshape&lt;/strong&gt; this long, we need to identify all &lt;strong&gt;stubs&lt;/strong&gt; which is easy enough in this case with 3 variables.&lt;/p&gt;

&lt;p&gt;[code language=”r”]&lt;br /&gt;
unab stubs: *90&lt;br /&gt;
reshape `stubs’, i(make model) j(year)&lt;br /&gt;
[/code]&lt;/p&gt;

&lt;p&gt;Unfortunately, variables not in the dataset for 1990 won’t enter stubs and you’ll be left with each instance of it (e.g., var91 var92 … var10) even after the reshape.&lt;/p&gt;

&lt;p&gt;The fix is to identify all variables in the dataset with any valid index value, but to ensure you end up with a vector of unique variable names. This is relatively straightforward with a combination of &lt;strong&gt;unab&lt;/strong&gt; and &lt;strong&gt;uniq&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The data I’m currently working with has denoted 61, 71, 81, and 91 with suffixes of 6, 7, 8, and 9 respectively. Some variables exist for 71 onwards, others only for 61. I therefore want to identify all the stubs with the following.&lt;/p&gt;

&lt;p&gt;[code language=”r”]&lt;/p&gt;

&lt;p&gt;unab stubs: *6 *7 *8 *9 // local list of all variables satisfying wildcard conditions&lt;/p&gt;

&lt;p&gt;loc all_vars “”&lt;br /&gt;
foreach x in &lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;stubs’ {  
loc x\_sub = substr(&quot;&lt;/code&gt;x’”, 1, length(“&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x’&quot;) – 1) // stub name is variable name less numeric suffix  
loc all\_vars &quot;&lt;/code&gt;all_vars’” “&lt;code class=&quot;language-plaintext highlighter-rouge&quot;&gt;x\_sub’ &quot; // ensure space follows &lt;/code&gt;x_sub’ to avoid smashing all variable names together when concatenating&lt;br /&gt;
}&lt;/p&gt;

&lt;p&gt;loc all_vars: list uniq all_vars // ensure stub names are not duplicated&lt;/p&gt;

&lt;p&gt;reshape long `all_vars’, i(district state) j(year)&lt;br /&gt;
[/code]&lt;/p&gt;

&lt;p&gt;Now things can go horrendously wrong if you have variables incorrectly ending in one of the stub suffixes. An easy fix is to ensure consistent endings, like renaming pertinent variables to *_6 instead of *6.&lt;/p&gt;</content>

      
      
      
      
      

      
        <author>
            <name>admin</name>
          
          
        </author>
      

      
        <category term="Computing" />
      
        <category term="Data" />
      
        <category term="Stata" />
      

      

      
        <summary type="html"></summary>
      

      
      
    </entry>
  
  
</feed>
